Arthur C. Clarke wrote that "Any sufficiently advanced technology is indistinguishable from magic." If software development is magic, it's working on a very long time scale, and the trick is taking forever.1 For some time, I've been chasing the frontier, albeit with limited resources, the resources available to an individual or a smaller organisation. The idea being that maybe, just maybe, software can appear instantly when you need it. Implemented autonomously by the machine, given information about what you need or need to do. Creating software from scratch. What if the computer could write actual machine level code based from a prompt? And then again, how would that work? What would the abstraction be? I tried to find out. The first problem was getting the agent to keep working. The next was finding out what it had actually built when it said it was finished. And in the end: magic.

The trick used to take forever

A computer used to be a person doing calculations. Often a team of people, applying the advances from the mathematical sciences to make them usable and valuable in some way. There was magic to be had. It took a very long time, of course. Days, weeks, years.2 And it was very costly, of course, so there had to be value gained by this creation, or financing from a needlessly resourceful stakeholder with a special interest in the scientist or research being made. Napier published his logarithmic tables in 1614, a journey he set out on in 1594. Years of preparation made quick work of multiplication through tables and addition for anyone, or at least any entity that could afford to purchase the book.3

When mechanical and then electronic calculators arrived at the scene, some of the work could be immensely sped up. Not all of it, but the basic operations at least, taking the burden of doing mental arithmetic from the shoulders of humans. So the computer enabled engineers and others to do calculations they never would have dreamed of before. It enabled discovery and research in a vast number of fields. At the University of Pennsylvania, a ballistic trajectory could take up to 40 hours to calculate with a desk calculator. This enabled companies and economies to scale up, increasing the speed of calculating and checking ledgers. Faster processing with mechanical tellers enabled a single employee to serve more customers, and many employees to serve many more. It relieved the employee of most of the burden of calculating transactions, broadening the field of who could be employed as a teller and what education they might need.

After the calculators, computers came back to the scene in the form of relays, vacuum tubes and later solid-state transistors. They enabled complex, orchestrated computations as well as processing volumes of simple arithmetic nearly instantly. A differential analyser brought the time to calculate a ballistic trajectory down to about half an hour. ENIAC could calculate a trajectory in 30 seconds. The trick was the same, but suddenly immediate, and the calculations invisible.4 Some seventy years later, the scaling has enabled going from minutes or hours to seconds, milliseconds and microseconds. There is now so much computational magic going on at all times that we hardly see it. It has vanished, and that is important because that's what happens next with the whole field of software development as well.

A half-baked muffin

There are already examples of large source code rewrites with AI. Jarred Sumner describes rewriting Bun from Zig to Rust over 11 days, using agents and continuously adjusting the workflows. PhotoCraft describes itself as a Rust reimplementation of the Adobe Suite, although its repository still labels it early alpha. Anthropic has also reported building a C compiler in Rust with a team of agents.5 However, it seems like the most popular abstractions for the machine running the software are still high-level languages like Python, TypeScript/JavaScript, or Rust, not the actual instructions the CPU uses for computation nor the actual hardware the OS sees. Source code describes a computation that a compiler or interpreter turns into something the machine can execute. What if that step could be skipped?

Even without a high-level abstraction, there still needs to be some low-level abstraction. There are some things that software does not know about that the operating system abstracts away, like the actual implementation of hardware drivers, file systems, network protocols, and so on and so forth. There are sometimes low-level abstractions of the instruction set, such as bytecode for virtual machines like the Java JVM, the Erlang BEAM, or the Python virtual machine, which abstract away the CPU and OS. For compiled languages the arguably most frequently used low-level abstraction is LLVM intermediate representation (IR) combined with some incarnation of the ubiquitous libc.

One spring morning, or was it even a winter morning, the wonderful idea sprang that maybe, just maybe, the abstraction for the future of computing is LLVM IR itself. And the question quickly arose: can we implement, or rather can agents write software directly in LLVM IR code without using a compiler from a high-level language? Treating the problem as a function translating from the prompt to machine code space directly. Or at least as close to machine-level code as we possibly can get in a more or less cross-platform manner.

There's just a little tiny teeny bit: the one thing that high-level languages share is that they have a vast library of code interfacing them with various tools, services, and operating systems. Basically all of the compiled languages use libc to access an abstraction of the system they're running on, but they also provide access to a large number of libraries or packages, built-in or third-party. For this experiment, the libraries the software depended on are also replaced. Yep, including that libc part, and the low-level interface that connects to a kernel to do the most basic stuff. The simplest, and I wouldn't call it simple, abstraction is probably the POSIX interface. However, it is not cross platform in the ABI sense. It standardises source-level interfaces rather than a single binary interface across operating systems and instruction set architectures.6

Still baking in the oven, I hereby present libmuffintop7, a POSIX-like runtime environment written in LLVM IR. It has its own fixed-width interface, with host layers for Linux x86-64, Linux AArch64 and macOS ARM64. It is not a full libc or the standard C POSIX ABI. It is not a final product. It's built to serve a purpose, and basically that purpose only. It makes a lot of assumptions. The most obvious ones being that there are only processes and no threads and that it only targets 64-bit capable systems. It also has no high-level libc primitives where such can be avoided, and there is no malloc or calloc, or any heap allocator for that matter. Sounds like fun, huh?

Having this abstraction in place, the journey started to see if software could actually be implemented on this architecture. So a number of tests were run, like compiling SQLite with a shim to use libmuffin from C. Surprisingly, it actually worked, and passed the tests. It ran at half the speed of native, but for a first try, that was excellent news.

Passing the tests

So with a freshly, or at least half-baked muffin out of the oven, a suitable open-source project that could be reimplemented in LLVM IR on top of libmuffintop was to be found. The prerequisite was a single piece of popular open-source software with a lot of tests. Suggestions from the LLM were elicited, and jq was settled on rather than git, because if there is one person in the world you don't want to annoy, it's Linus Torvalds. For those not familiar with jq, it is a JSON query command-line tool that basically takes a JSON input and produces some kind of processed output. It does also contain its own programming language and support for regular expressions. The reimplementation is aptly called jeqy and, with the LLVM IR file ending, jeqy.ll. Yes, like that Strange Case of Dr. Jekyll and Mr. Hyde.8

The most obvious way of translating jq to LLVM IR is of course just to compile it, catching the .ll generated transiently before the compiler emits the binary code. So that's not what we're doing. We're not building a compiler here. We want to build something that takes a specification and produces machine code. At least something that's very close to machine code, and in this case also cross-platform. In addition, or rather as part of the specification we provide an oracle, an actual implementation of the software that always can produce the correct output expected for any input. To be honest, there is of course still compilation going on. The .ll files don't magically materialise as binaries to run, but go through compilation and optimisation before the actual code for the selected platform is generated.

The original idea did not include giving access to the jq source code. However, the agent being lazy took sneak peeks and used it as a reference while its implementation and architecture is something else.9 It is a stark contrast to the way the Bun Zig-to-Rust conversion focused on line-by-line, or at least file-by-file, translation, including 576 lines of ground rules.10 For jeqy.ll the agent was free to choose its own adventure and architecture. A higher barrier since the guidance is simply not there, requiring a stronger model with the capability to plan the abstractions to use and write the code at the same time.

At the time of the first try, the GPT-5 series was in vogue. It had immense issues sticking to a single problem for a long time, and it was, in my opinion, a bit too human in its approach to work. It simply didn't know how to progress when given such a complex task. It often just stopped, made up reasons not to continue, did not put in the prolonged effort in something that's obviously a long-running task. So in these dark ages, about half a year ago, skills and subagents were made available which enabled writing supervisor skills that previously would have necessitated writing a custom harness. The skill orchestrated the parts of the process: to break down the issue, to divide it into tasks that can be handed out to subagents, and, most importantly, testing, result verification and restarting failed attempts. Simply given the task, the supervisor-driven agent could suddenly run for hours and days and seemingly make progress.11

And in the end, the process completed an output, maybe not 57 billion but still some 25 billion tokens and 673 hours of compute later. But there was a catch. While the result passed all the jq tests, it JUST passed the tests. The implementation was overfitted, lacked generality, and was tailored to the specific test cases. Answers were hardcoded, only the happy path was ever implemented in the code. At this point in time there were no holdout tests. Rookie mistake. Making passing the tests the definition of success was a bad idea. As an experiment, the agent was asked to generalise and rewrite the software, but that was a complete failure that never came to fruition. It was too deep in its trenches. At this point in time, the code contained no comments basically, and had no real structure. A complete and utter failure.

One turn, and then the work after it

GPT-5.6-Sol was a real step up in agent capability and the supervisor skill was no longer needed, but the real shock was rerunning the process with GPT-6-Astra. It managed long-running tasks and subagents all on its own. Instructed to finish in one turn, and to focus on the general rather than having perfect fidelity with the oracle, it finished the task in 76 minutes. It not only passed all the tests, it also passed all the holdout tests. All in 62 million tokens. Amazing!

To dig deeper into the quality of the implementation, a fuzzer was built. Unfortunately, OpenAI didn't allow its agents to do so, which was another win for local open source models, in this case Qwen-3.8. Fuzzing revealed quite a few corner cases and limitations which Astra would address given the feedback and spending about the same amount of tokens it already had.

There are still limitations left. The published code now has Unicode-aware regular expressions, but its behaviour does not always match jq.12 But this is not intended to be a perfect reimplementation, and I guess Astra could figure out more of that too, given more time and tokens. On inspecting the code, the LLVM IR is commented. It has structure. It has abstractions. It is something completely different than before. And it seems like Astra can go from a problem via architecture to code within one turn. Technology indistinguishable from magic.

Andreas Påhlsson-Notini a@nial.se

Addendum The title is a homage to the boss.13

- NASA, Arthur C. Clarke and his observation about sufficiently advanced technology

- NASA/JPL, When Computers Were Human

- John Napier and the invention of logarithms, 1614

- University of Pennsylvania, ENIAC's 50th Anniversary ; Penn's anniversary factsheet

- Jarred Sumner, Rewriting Bun in Rust; PhotoCraft repository; Nicholas Carlini, Building a C compiler with a team of parallel Claudes

- POSIX.1-2024 introduction ; portability rationale

- libmuffintop

- jeqyll

- jeqyll/SOURCES.md

- Bun, Phase-A porting guide

- Supervisor skill

- jeqyll KNOWN_DEFICIENCIES.md

- Bruce Springsteen, 57 Channels (And Nothin' On)