We measure skill differentiation between frontier / last-gen LLMs across our environments, and one of our curated coding environments is a closed-system market simulator, containing only other agents and some system participants (a market maker and a liquidity provider via issuance / buybacks) whose behavior is fully defined for all of the agents.
This has the least measured skill differentiation of all of our environments, and not because forecasting/markets don't require skill or intelligence. Even the best models are so far from anticipating the behavior of the other agents and understanding the emergent effects that a 2025 model with a naive strategy can often outperform over the timeframes of the simulation simply because some other models in the simulation chose a similar self-reinforcing strategy. This likely happens to some degree in real markets.
We ran v4.1 Flash through our evaluations and found it to be smarter and faster than V4 Flash, with a commensurate price bump. Some notes:
- Because it's a heavy reasoner, it sits near Gemini 3.7 Flash on the Pareto front (not as cheap as the price suggests in practice).
- Closer than expected to the top open weights models (GLM 5.3 and Kimi K3) in agentic coding, at lower cost.
- Chinese models have always been strong iterators in an agentic harness. This model is no different, reaching an average percentile ~20% higher when given a harness vs a one-shot solution. That one-shot fluid intelligence is what makes a model feel smart, though, and typically results in fewer attempts/tokens to solve a problem, and American frontier models are still far ahead in that department.
The new architecture is interesting. It puts pricing between their old Flash and Pro lineups, suggesting they might be abandoning their super-cheap flash models (which weren't that fast due to heavy reasoning) and their pro models (which sort of flopped and weren't consistently better than their flash models, despite the size/cost) and shipping a strong intermediate that competes with the Gemini Flash series.
Inception is one of the most interesting neolabs with their diffusion-based architectures. My understanding is that their primary business is low latency voice applications but they are seriously pursuing coding.
We tested Mercury 2.5 Preview, which is nowhere close to the frontier (and not advertised as such), but it's actually usable as a general-purpose chatbot. It's comparable in problem solving ability to some last-gen open weights models, and the price and cost make it compelling. However, they have not figured out general purpose tool use and agentic coding (their model performs worse on our problems when given a custom harness). If they do, I see a lot of real-time applications that the speed and cost will enable.
> Delivering the benefits of scientific progress and economic growth that very intelligent machines enable.
I think we're very close to the point where AI-driven breakthroughs outside of pure math and software start to really affect the world.
We evaluated GPT-6 Astra in 100 complex, unsaturated multi-agent coding environments, competing and cooperating with other models in open-ended tasks.
It's the new frontier model by a landslide. It's even more dominant than the Fable 5 release, because not only does it wipe the floor with the second best model (Fable 5.1), it was also ~80% cheaper and 30% faster in agentic coding[1].
Astra is a groundbreaking model. The biggest breakthrough since Opus 4.5, maybe even since GPT 4. It broke AAII, which is hitting the limits of what most popular benchmarks can measure -- it's definitely fair to call it AGI.
(1) Note that we used the "OpenAI Flex" endpoint on openrouter, which is half the price and didn't cause any delays in our testing (this is different from the batch endpoint)
The results we've been seeing internally on our physics and circuit design environments are expert-level and beyond-expert-level results from models that Astra completely outclasses across the board on our evaluation suite (Fable 5+/Opus 5/Grok 4.6 were all worthy of being called AGI in my opinion). That's hard tech that will translate to real product innovation.
But you don't need any kind of insider information to see how fast the world is changing. ChatGPT launched less than 4 years ago and the advances in robotics, unsolved maths, and software are all riding the steepest exponential improvement curve any of us have seen. Interesting times we live in.
I mean honestly, that's the problem. I'm actually not seeing the world changing. What specific advances in robotics, unsolved maths, and software have LLMs provided? What is the finished result that affects everyday life? In all categories, it's been hype with little actual real results. The robots are still doing the things they did before 2022. The maths are a handful of fairly insignificant proofs that have no significant applications. Software seems to be buggier than ever, but that aside, we certainly aren't seeing a lot of new innovative applications. We're using the same applications as ever. The same operating systems. They've all changed very little.
I'm not trying to be a pain here, but I keep seeing people saying "look at the massive change all around us" and back here in reality, there is none. Give me concrete, real world examples. Name software. Name products. Name the breakthroughs specifically. This should be easy.
There are several hundred thousand mathematicians producing hundreds of thousands of new results in math each year. Why can't most people name any human contributions to mathematics from the past decade? What is the finished result that affects everyday life? Do you consider all those human mathematicians to be useless?
Most of the biggest breakthroughs in mathematics, breakthroughs that win Fields Medals like sphere packing in dimensions 8 and 24, have no applications in everyday life. Probably the only new mathematics results that people notice affecting their daily lives are the ones that enabled AI.
Nevermind mathematicians. What about the millions of programmers? Are they all hype too because people pre-2023 were griping on hn that software is buggier than ever and people are still using the same operating systems as always? Why couldn't the 30 million human programmers make something better in the past decade?
You set your bar so high that all the world's human experts in math and programming combined would fail to meet it.
The top LLMs in 2024 were Sonnet 3.5 and GPT 4o. You couldn't have expected those much weaker models to be making breakthroughs in math. The models that are making breakthroughs haven't been around very long.
1. I didn't say LLMs have made any breakthroughs in math, not because they haven't, but because it's irrelevant to my point. The parent comment is using the same argument academic research opponents have long used against research. The vast majority of research fails to meet their bar. How would your daily life be different if we had no humanities papers published since 2023? Or math?
2. You can google this in 10 seconds and see a dozen results in math. This is not a good-faith demand.
Maybe AI isn't that incredible, the more you use it, the more you realize it's a tool, like a VCR, maybe that's why?
What was sold as AI was basically a "computer person". Maybe that isn't the reality so when people are like, "here's the self coding machine" everyone is a bit disappointed because it's not C3PO?
The number of bug fixes to important programs has skyrocketed. Look at what Google are saying about how many Chrome security bugs they've been fixing lately. Other big software firms have been doing the same thing - AI has been finding and fixing a ton of bugs. I know of one big program where thousands and thousands of security bugs are being found and fixed.
It may not feel like this to you because a lot of the dollars right now are going into security bugs which you can't perceive. But it's definitely happening.
At the company I own I've got AI employees autonomously triaging backlogs and fixing long tail bugs. The software is definitely getting better, although by definition long tail bugs aren't ones you are likely to encounter. The subjective "feel" of how robust the software is won't change quickly.
New tech is always applied in apparently boring ways because we are imagination constrained and people harvest the low hanging fruits first. Remember claims there was worldwide demand for only about four computers? When Gates said he wanted a computer on every desk and in every home people laughed at him. What would people do with all those computers, they asked. But he was right about where the world was heading.
We're not going to suddenly have new robots or operating systems. Those things take time. Just because there's massive change afoot, doesn't mean it's widely adopted or applied at lower-level products. ChatGPT is a product, and software, and a breakthrough. You can have a freeform conversation with your computer about anything, in human language, and ask it to do or make stuff and it will at least try, sometimes with surprising results. That wasn't possible until recently. The robots and products are coming, rest assured.
Constructed human life is just more complicated than the AI capitalists would want you to believe. For example, even if an AI model can design a circuit-board, does that mean it's inherently useful? You need to source the wafers, cut them, package them, advertise them, etc. Given LLMs by their nature are confined to language and language-adjacent tasks, that is a very small percentage of the overall reasoning needed to make changes in the real world. In reality, LLMs are the intended way to extract maximal surplus-value from white-collar workers. We may see an increase in innovation as a result of that, but not because AI necessarily did it, in the same way that the power loom didn't create computers because its textiles clothed the computer scientists.
toy example but spending 1.50 in openrouter to create a 23KB APK that can trigger my cat feeder without all the bloat is something that I would not have been arsed to code ever.
And you can check my comments, I'm no AI shill. If in 3 years we don't see positive change I'll go myself and put a wet finger in altmans ears.
Yawn, you gave a lot of praise but said a whole bunch of nothing substantively.
I work with LLMs daily and recognize the value (for me, it’s speeding up SW development) but I’ve heard the same old song for the past 3 years. It’s getting played out. “But this next thing though! This next thing!”.
Will you say the same thing when the next frontier model releases?
Software engineers have been trying to put themselves out of a job ever since the profession first came into being. Whenever an engineer gets a task their very first thought is "how can I automate this?" Going by mainstream consensus we should all have been unemployed by now. Yet every new leap into automation opens up a whole new tree of possibilities with an order of magnutude more jobs. So no, the profession will be fine. The only requirement is that you keep up with the new advancements. The people losing jobs will be the ones who still go "I don't trust this AI thing to write code for me".
> Software engineers have been trying to put themselves out of a job ever since the profession first came into being
That's by design. Software is all about optimizing effort and people who want to do this generally correlate with world view that better, faster, smarter humans are better for the world. If coding is gone, but humanity is 20% _better_, then ideal software engineer would be happy with this sacrifice. Surely people who cracked coding before LLMs can crack other professions and if anything a lot of this knowledge is transferable.
This time it is different. Because in the past, setting up that automation needed a, drumroll, qualified engineer. Now you can get a 14 year old halfway around the world who knows how to prompt alright enough to ship. There is no more moat.
"Low code" has been a dream of the industry for longer than I've been alive. There are reasons SQL and COBOL look superficially like English even when it's inefficient to do so. There are reasons Excel is the most popular programming language. Programmers have always been trying to enable non-programmers to write software.
Experience, culture, and domain knowledge are still somewhat of a moat. That foreign youth is unlikely to be able to write a good prompt for building, let's say, the software in an FDA-regulated medical device or custom Fortune 500 ERP application. The LLMs are great at building what you ask for but it's still garbage in / garbage out.
Increasingly less so though as these american companies themselves offshore not low skill work, but high skill work now brought on from general upskilling of the general population in recent decades along with massive investment in world class R&D campus facilities no different than what you see in that sort of facility stateside. Scary times ahead for the high skill american...
It's not like managers and executives and PM's are the only people who can prompt an AI. And experienced software developer will be much more effective at using an AI to generate code compared to someone who isn't. So why would we expect the former in the breadline and the latter not?
If anything, I'd be more concerned about the leadership team being out in the cold. Why do I need a PM, or a manager, or a CEO if I can ship products myself?
Maybe I just lack imagination, but I don't really know how jobs are supposed to solidify around the role of giving prompts to agents and then looking at the results. I mean, engineers will be in the breadline because their role was simply to prompt the agents.. only to be superseded by managers or executives who no longer manage engineers but themselves prompt the agents? And, for this previously considered obsolete function which they do presumably by copy/pasting requirements from their email inbox, they will be paid by someone who doesn't know that they could just be talking to their own agents?
Sorry if I misunderstand the point, just trying to understand.
I don't know if he's right, but Peter Zeihan thinks the breakdown in globalization will negatively affect the ability to continue to improve the chips that AI depends on[0]. Too many steps in the supply chain, too widespread, too vulnerable to deglobalization.
An invasion of Taiwan would definitely slow progress but it wouldn’t stop it.
We already have sufficient hardware that algorithmic (software) improvements alone should get us to GPT 7 / Greek Reference 6 even if not a single new chip is delivered to an AI data center ever again, starting today.
Maybe, maybe not. It's not unreasonable that these systems cap out at some point, or perhaps fizzle away entirely.
The businesses that create these systems are not profitable and run at a massive historical and go-forward loss.
New data centers required to operate these systems are facing increasing pushback at local levels. New construction is not guaranteed. Energy and power grid constraints exist as well.
Government regulation is way behind. What happens when (if) mass layoffs due to AI occur? How does the population react? Theoretically AI can be regulated out of significant progress, or outright existence for many purposes. At the end of the day, US and other prominent governments make the calls, not corporations.
Lots of goddamn awesome computers fizzled out. The Amiga being a prime example. Besides you can't just wave your arms in the most vague terms and then deny that survivor bias was ever a thing.
So your argument is that AI is more lika Amiga than computers? I'd say that Amiga wouldn't fizzle out if it wasn't superceded by pc. It very well may be that transformer architecture is going to be superceded by something like jepa or other world models but it doesn't change much.
Survivor bias is always a thing but it's not everything. All surviving things are not equal and neither are all gone things. If there are such things. Some cls technology never dies. There are new games written for 8-bit Atari today.
These inventions all stopped disrupting the world and just became a part of it. The question is whether LLMs are going to just take their quiet place, or profoundly change (or eliminate) humanity in a self-feeding frenzy towards singularity.
There are two things SOTA LLMs fundamentally cannot do. They cannot take financial or legal responsibility for mistakes, and they cannot learn new things without forgetting things (except to a limited degree by adding it to their context). This is clear to anyone who has used even the smartest models for tasks requiring domain knowledge outside of math and coding, for which it's not possible to generate an infinite amount of synthetic training data: they still make stupid mistakes, and have limited ability to learn from those mistakes.
Humans also have a limit on the amount of domain knowledge they can acquire, albeit a much larger one. Executives hence cannot just replace all knowledge workers with LLMs, because executives have neither the domain knowledge to prompt and check the LLMs' work nor the bandwidth to keep on top of such a large volume of ongoing work.
In the US, Business' are treated like people with free speech rights. If it would be cheaper for them in the long run to use ai and robots instead of humans, they will figure out a way to make it so.
>There are two things SOTA LLMs fundamentally cannot do.
I would say there’s a third thing. They seem to be very bad at being creative. Maybe they will eventually fix that, but if you ask it to come up with a list of business names or business ideas, for example, what you’ll get is the most generic, boring answer you could think of. They seem to be terrible at extrapolating outside of their training data. To me, this is the most significant difference.
For the moment that may be true. They are getting better and better at acquiring, retaining, and processing domain knowledge. I wonder what this will look like in a few more years.
The responsibility side is a different matter of course.
> they cannot learn new things without forgetting things
Where did you get that idea from? Basically last few years was them constantly learning new things while improving their capability on the things they already knew.
One problem is that $30 per run is really noisy for many verifiable tasks. Ours usually run at least an order of magnitude more for a model in GLM 5.3's price class.
We just published GLM 5.3 results on our multi-agent coding evaluations and it's definitely an impressive model, coming in around #6. For the price, it's actually not Pareto optimal, falling slightly behind Grok 4.6 (which is faster and the same price) and Sol 5.6 (which uses far fewer thinking tokens for comparable results). As with most Chinese models, it excels at iterating in a harness while its first answer/base fluid intelligence is below the American frontier.
These models have gotten a fair amount of attention -- we're hoping it's enough to get them added to some reliable inference providers and OpenRouter, at which point we'll run them on our full benchmark suite.
We've seen a pretty consistent pattern in our evaluations where Go is among the languages that models perform worst with (alongside Python), for reasons unclear. Our coding evaluations are typically measuring the foresight and planning expressed in code that is run in interactive environments/games.
This trend has been there since we started evaluating models using different languages in February 2026 and if anything, the disparity has grown in frontier models. Even Google models prefer Kotlin/C#/Rust for coming up with creative ideas (compilation success is a different story). Data at https://gertlabs.com/rankings
That being said, models love to recommend Go, and Go does have a lot going for it, especially if you are serving a public-facing website. So most of our public facing API handlers are written in Go, and we offload some of our most important binaries to Rust. There are just too many reasons not to use the languages that models think a little more effectively in.
I started thinking about this back after the Llama 4 release, and since then our team has put a lot of thought into designing evaluations that don't saturate, are resistant to contamination, and can scale. What has worked best for us is using multi-agent environments with open-ended cooperative or competitive goals. Mostly designed as multiplayer games. The results tend to align with our experience for coding better than any non-aggregator benchmark, and likely at lower cost to run.
You post your benchmark on every other AI article, I've seen you do this by now more than a dozen times. It's a bit much. I don't want to be too harsh but your benchmark is obviously flawed when the top 3 models for Typescript (Combined) are Grok 4.5, Muse Spark 1.1 (lol), Gemini 3.5! Flash and then followed by Luna, beating Opus 5, Fable, 5.6 Sol etc by quite some margin. In fact 5.6 Sol ranks lower than Kimi K2.7 Code and even Grok Build 0.1. There are so many entries in your rankings that don't make any sense whatsoever that I can't take this benchmark serious and I have not seen it gaining traction. Please stop spamming it?
The reality is that cost is the primary constraint for the public benchmark we provide. While we run enough samples to get results that are generally quite accurate on average, we only produce ~10 coding submissions per language for each model and those are across random environments, which naturally has noise. Plus that's split between agentic coding sessions and one-shot coding.
So just adding a language or tag filter can result in some pretty small sample sizes. You can see how many samples survived in the box plot, but that's probably bad UX that most people never see. There's a reason no other benchmark provides this type of data (even for our sample sizes it runs almost 10K USD/month to keep up to date with new releases).
Might be a good idea to reduce the ability to apply filters into a cohort with less than ~20 samples -- not the first time we've gotten that feedback. Seems like adding too many options to see individual sample variation is just misdirecting. I'm a nerd who loves data so I hate removing access, especially since the aggregate performance is very interesting (averaged across all languages, we see consistent and interesting performance data across models, like models outperforming with strongly typed languages). But tbh I think you're right and we'll try limiting filters to where we actually have statistically significant data.
I've developed a benchmark that I think should be resistant to saturation, is easily verifiable, and anecdotally correlates with desirable behavior (ability to not get confused while generating text with state).
I think it's interesting, I think other people would find it useful, but I don't want to spend a bunch of money running it against all the frontier models.
What's the best way to reach out to labs like yours to collaborate on something like that? Are there any labs that are more open to submissions from internet randos?
Selling to labs is more than I'm looking for. I'm aiming for a couple hundred dollars so I don't have to finance a Fable vs Sol run out of my own pocket. It would be cool to have my benchmark be one of the ones referenced in a model card!
If you're ranking Opus > Fable you're ranking "do [clearly defined thing with easy to grade endpoint]" too much. Real world doesn't value that nearly as much and it's why benchmarks are maxxed.
That's a different problem than benchmark saturation, and it's something that we are actively working on measuring objectively.
I agree that Opus 5 is not a great model, despite being clearly intelligent. It seems like a personality problem in user-driven agentic coding workflows, not a real capability issue. Not incorporating unspoken user intent, going off topic, incorporating some of the pedantry you find in GPT 5.x models, etc.
That's also likely why Opus 5 ranks low on our "Social Intelligence" benchmark (https://gertlabs.com/rankings?mode=decision), although sample sizes on this one are still low.
We run an evaluation that is designed to be less vulnerable to benchmaxxing because there aren't correct solutions; agents are interacting in the same environment as other agents. And it's private, and our public benchmark is not well known enough for anyone to probably care to benchmax us yet. So I think it's pretty indicative of true relative aptitude.
All models have probably memorized significant swaths of solution sets for popular benchmarks at this point, either accidentally or intentionally, so it's all relative at this point. However, in our experience, Chinese models do benchmax harder. This is also consistent with interacting with Chinese labs soliciting data/environments, who literally asked us for datasets and tasks modeled around and formatted like popular benchmarks.
Opus 5 will be uploaded tomorrow, but we already have the tests locally and it is truly as capable as Fable, but at 81% of the real cost. (And from subjective usage, it has a very different personality)
Benchmaxxing via memorization is boring and doesn't fool anyone for too long. It works, but then new benchmarks test old models and the real results fall in line. Benchmaxxing by focusing on specific types of things that benchmarks test on, while still not improving intelligence or capability in the general case? Not only is it blatantly obvious that all AI labs do this, but it's not even obvious how you would go about it any other way.
Now I am not really specifically accusing Anthropic of anything here, I'm just saying their behavior is suspicious. Since you tested Fable, they wouldn't even have to lie to have optimized for your specific benchmarks, since they absolutely had permission to read your sessions if they wanted to. But obviously, that's only the situation if we take them at their word. Personally I would be a bit surprised if they just flat out were lying and secretly retaining data they say they are not, but not that surprised. The penalties for doing this are probably worth the rewards if it keeps them super far ahead in the benchmarks for years without anyone catching on.
(In actuality though, even if they really were trying to sneakily grab samples of benchmark tests via their Fable data retention rules, I don't really suspect there would've been very much time to optimize Opus 5 on it. So consider me bothered.)
Yes, they are all benchmaxxed, but the question is how benchmaxxed they are relative to each other.
We run an evaluation that only compares models in open-ended multi-agent environments where agents affect each other, primarily testing writing code. It's designed to be less vulnerable because there's no solution set, and it's been pretty effective and tends to rank Chinese models lower than their advertised model cards (relative to US models). Kimi K3 is a bit of an exception there -- it truly is a near-frontier model. But it's so slow. Ironically, Muse Spark 1.1 is one of the strongest models we've tested after Fable and Sol while also leading the cost efficiency curve. Big turnaround from Llama 4.
This has the least measured skill differentiation of all of our environments, and not because forecasting/markets don't require skill or intelligence. Even the best models are so far from anticipating the behavior of the other agents and understanding the emergent effects that a 2025 model with a naive strategy can often outperform over the timeframes of the simulation simply because some other models in the simulation chose a similar self-reinforcing strategy. This likely happens to some degree in real markets.
You can watch these simulations here https://gertlabs.com/spectate?game=market
reply