How I view LLMs as Sept 2026

I’ve spent over $30k in tokens since the start of the year across Codex, Claude, and GLM.

Here’s how I view LLMs as of September 2026. Two categories:

Instructions-following models:

- Like Luna and Sonnet.

- I used to call these “dumb”. But that’s the wrong way to think about them.

- They just cannot make good “common sense” decisions.

- They’re workhorses for pennies. Luna especially, is hard to spend a lot on.

- If you make the core decisions up front, coding is auto-complete.

- Frontiers like Astra and Fable.

- These are higher “common sense” models. These models (try to) understand the intent behind the question.

- Solving problems is about making the right tradeoff - these models tend to make the right tradeoff more often than the other models.

- These models know how to orchestrate instruction-following models prescriptively - in a way humans get too lazy to.

In my experience, today’s models are far from making the right human-like trade-offs.

However, this is where I believe the path to AGI is.

That is, if a model can make the tradeoffs a human would.