How many people actually read the full post? It builds up this amazing underdog story where all the benchmarks are taken as victories, and then only late in the post and section 6 is it finally revealed that the only way they won was to fine tune directly on the benchmark.
1. It’s hard to trust a 2026 paper that’s showing results for such old models.
2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on skills that generalize to make them experts at unrelated tasks.
3. The study sounds like proving humans will never fly because they don’t have wings. In reality, humans do fly, and Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.
Not sure how that vague truism applies to this paper.
Lots of papers have great results that don’t depend on the latest models.
However in this case it’s problematic:
- They specifically make claims about the state of “current LLMs”. o3 is not representative of this.
- They ask are LLMs capable of X and arrive at a negative result.
If their claim was LLM’s can write coherent sentences, and their conclusion was positive, then there would be no issue using old models because the end result would be factual.
However, when you have a negative result that makes a claim about the current state of all LLMs and the ones you were using are not current, by definition it draws the whole conclusion into question.
Maybe practically it doesn’t matter? Perhaps AGI is not the model but the model plus everything it’s got access to. If we’re modelling intelligence in the way we seem to have to to have any coherent definition of AGI, it seems to me <model + everything it can access> is always going to be more “intelligent” than <model> alone.
My points is more that, while we have a strong intuition about where, as an entity, a human's boundaries are (i.e. where the person begins and ends), philosophically it' not immediately obvious that the analogy applies to the a model in the same way. Why should that be the line drawn that says this is the "thing" and this other stuff is external to the thing? It feels somewhat arbitrary.
Of course this is a difficult question with humans too, hence my reliance on intuition above. We don't have the same cultural/biological framework to fall back on with AI.
I think all this debate about whether an LLM can write (or download) a chess engine is sort of missing the point. For basically any economically valuable work there is no equivalent of a chess engine for it. If there were we wouldn't need humans or AI to begin with.
If the goal is merely to "win at chess", then yes, an LLM using stockfish is better than any human alone at performing the task. When you are talking about what AI agents are capable of doing, there is no such thing as "cheating". They are as capable as the tools they can use effectively. The entire history of human civilization was driven by effectively using tools to achieve goals.
so prove it! get a public repo out there, have it play against some open source engines
also I think the operative letter in AGI is the G - and if the G is short for 'variably competent savant-like hyperfocus on certain kinds of software coding and not any other general skill' then its not really G at all, is it?
Declarative knowledge is not the same as procedural knowledge. You can read as many chess tutorials, strategy documentation and game archives as you like, they won't make you good at chess until you actually start practicing chess.
The issue with the models isn't that they play a bad game, but that they persist in making illegal moves. An average intelligent human can be told the rules of chess and then play chess, badly, within the rules.
Sure, but the LLM is free to construct a representation of the chess board and update it as it goes along. It is not in any way banned from using a virtual board, or whatever representation of game state it pleases.
AFAIK, current models will still sometimes make illegal moves even if given the entire game state (e.g. in FEN notation), so it is not purely an issue with the models’ ability to keep track of sequences of moves.
My point was you are misunderstanding G, or at least applying it erroneously here. Being good at chess is not a generalization of any other body of knowledge, it is a rigorous set of rules. The only way to be good at chess is to practice chess, or to apply deep calculations. The latter is the model writing code.
The illegal move aspect has more to do with a failure of online/in-context learning, which would support your point. I tend to think it is a byproduct of reasoning in language, which newer architectures would fix, but we shall see.
chess is not just a rigorous set of rules, it is rules as foundation with layers of strategy on top. and so is, for example, scientific methodology or chemical interactions or virtually everything else under-the-sun that comprises human knowledge
knowledge for chess is derived from memorizing strategies that have been well-defined for decades paired with in-game reasoning processes. this is not at all different from any other body of knowledge. Noble gases, laws of thermodynamics, organic chemistry just to name a few - these are all 'strategies' that define observed phenomena, analytical frameworks that trace a logical, rational set of interactions and which can predict the next
for an AGI, all of this should be a cakewalk, trained as it were to surpass human capability in any and every domain [0] (thus the G for 'general' and not 'N' for 'narrow' [1]). it should be a natural at everything, infinitely adaptable on-the-fly. the whole point of AGI is that it surpasses human capabilities even at our frontiers and bleeding edge (unless you're private enterprise and you've moved the goalposts for industry [2])
currently, it's only AGI-seeming if it gets benchmaxxed enough. otherwise it sucks at what it does and then is only barely competent at tasks if paired with enough skills and tests to make it more diligent at its work. this makes sense to me - for any probabilistically trained tool, even one that you post-train and fill with nothing but the best-quality evidence, the ultimate result is the lowest-common-denominator output for your sample set. there's no natural reasoning the AI does itself to make itself better at what it does - it's all human curation and categorization of sources ingested paired with RLHF post-training that we can get the mediocre-at-chess-at-best results that we see now and the benchmaxxed scores against whatever arbitrary and pre-defined measure
that's not AGI by any classical definition. that's a cool, useful, and powerful tool that makes our lives easier, much like a hammer, nail, and studs make mounting a picture frame easier than if we only had our hands alone
I'm not. AGI is almost necessarily closer to ASI than it is to human intelligence by definition. it's become pretty obvious that even what seems like irrelevant domain knowledge has utility applied to other domains - that's why we're pursuing general-use models
presumably, an 'AGI' that is generally as good as a really good human at every task under-the-sun will already be much better than most humans at the task because it can incorporate cross-domain knowledge and apply it in a reasonable fashion. it's like the parable of Newton and the apple - the domain knowledge that an apple falls according to certain rules observed through historic experience igniting the creative spark that led to universal gravitation
> presumably, an 'AGI' that is generally as good as a really good human at every task under-the-sun will already be much better than most humans at the task because it can incorporate cross-domain knowledge and apply it in a reasonable fashion.
I disagree with this definition of AGI, and I disagree that chess skills significantly benefit from generalizing non-chess knowledge, outside of computing moves probabilistically.
AGI has historically been defined as human level or better, with generality to new domains. I think blurring it with ASI makes the terminology confusing to use.
Chess is learned rules and the ability to apply those rules. Strategy as a whole is applying a set of rules to circumstances, that's how it is taught: "here are examples of circumstances and actions, try to pattern match to future circumstance and apply commensurate action."
If you make the point that chess is a large part of the training data, or that LLMs are unable to learn chess well, I'll accept that as refuting that LLMs are AGI, but these other points I disagree with.
> People who are good at it rely more on experience and deep domain expertise
People are good are 1900 or 2100 above and the top ones who spend decades in the field i.e. deep expertise are well in the 2200-2700 range.
A 1100 player is none of these things, they are purely relying on strategic reasoning there is a good chance they cannot name a single opening or articulate clearly why a move was appropriate. 1100 is quite low bar.
1100 at online speed chess or something, could be. I'm not that deep in the chess world but everyone I know that can make 1100 in official rating can name a dozen openings and most of the known tactics, and is pretty good at applying at least one opening.
1100 lichess/chess.com does not represent real elo. I'm around 1400 online, I would still be unranked in the real world. The fact that I easily beat any model publicly available is not a great look for AGI.
If the goal for buyers of AI is “replace this knowledge worker”, how much does it matter that the model in a simple loop can’t do it, but the model with a strong general purpose harness and a little time to gather resources and knowledge to augment the harness going forward, plus tool calls, plus custom built tools, etc, can replace the knowledge worker?
Probably the only thing saving many jobs from being replaced right now is that it’s hard to have a verification of correctness in the loop, so the agent can’t hill climb very easily.
Tests are often conducted under restricted conditions. For example, elementary school students aren't given calculators in math class, or during an interview, you are asked what encapsulation is and aren't allowed to use Google. The chess test effectively demonstrates the reasoning capabilities of an LLM without relying on brute force, because a human is incapable of calculating trillions of combinations yet plays chess successfully. This test is necessary because many complex problems cannot be solved by brute force, such as managing a business or playing Heroes 3. Therefore, we can make the assumption that if an LLM can play chess at a grandmaster level without brute force, it means it will be able to command an army or manage production.
People might care about this for chess, but no one really cares if an LLM can command an army or manage production of a business without any tools. If it can do those tasks reliably when given access to tools (including any tools it autonomously creates for itself), then that's more than sufficient. No one cares if an LLM is doing reasoning the way humans do it, as long as it can get the job done.
The assumption is that if an LLM is incapable of playing chess—a game with a relatively small number of pieces, clear and simple rules, and perfect information—even after reading a hundred thousand books on chess, then it is fundamentally incapable of managing an army or a factory. This is because those scenarios involve more 'pieces,' incomplete and fuzzy information, and implicit rules that need to be deduced independently. It doesn't matter whether it has tools or not. It's simply that running tests with chess is cheap, whereas testing with an army or writing a browser from scratch is quite time-consuming and expensive.
This is likely still an LLM (in the purest definition of a language model with relatively many parameters) since the inputs are natural language, just not a generative LLM as the output is something other than more language.
It is frontier in the sense it is exploring an unexplored domain. I do agree on questioning the comparatives though. Speed/cost is indeed relevant for problems that can be framed as structured decisions only. The question is, would defining a structured decision model be a structured decision model itself? This would significantly increase the application domain.
How is this not a frontier model? It's bleeding edge in its own niche. It's not a frontier LLM; however, applicable to many of the things people use LLMs for.
It's nothing like a traditional LLM and so should not be compared to one. It's a heavily constrained, tiny model that can only produce a probability score or a yes/no answer over pre-defined selections. It has no long-context capacity.
I mean, imagine comparing this thing to Astra, it's hilarious. They don't even tell you what the max input size is, and they only allow 10 possible answers to choose from for the Choice mode. It's probably like a 1billion param model. They say it's "not small", but there's zero reason to believe that.
I suspect someone will be able to recreate this within a week by piecing together open-weight models.
> It's nothing like a traditional LLM and so should not be compared to one.
Frontier LLMs are expensive jack of all trades. You can absolutely compare them to purpose-built tools on any domain they touch. Engineering is all about assessing tradeoffs.
His claim was that the title is misleading, not sure how it's relevant to that claim that you use "string models" (full LLMs).
The original title before it changed less than an hour ago was:
"Jev: New frontier model 40-400x cheaper and 20-200x faster"
I'm going to agree that was misleading.
And on the second point:
>>Also "can't hallucinate" seems wrong? Sure, it can't emit an invalid type, but it can still emit a completely wrong valid value.
>that is likely true of all ML! perhaps we could debate semantics, but I don't think it's fair to say a random forest "hallucinates" in the way LLMs do"
Also going to disagree here, and I don't think it's semantics.
No. Your launch post puts “0%” on a hallucination chart, then explains that the number comes from guaranteed schema matching.
You’ve already agreed that this doesn’t establish correctness. An approve for an unauthorized action still meets the schema guarantee.
That’s why I find the messaging misleading. You’re acknowledging the limitations in these replies while defending the broader reliability pitch.
Even granting that each answer is calibrated individually, that doesn’t establish calibration of the decision that combines them.
Sure, I can threshold a composite score, but there may be many wrong answers with the same score. An unauthorized action doesn’t become acceptable because it scores highly on the other dimensions.
I still have to define the constraints and test which wrong actions get through the complete workflow on my own data. That’s a substantial part of the work being pushed back onto the developer.
"Hallucination and type-safety are intrinsically related"
I'm not entirely sure why we're conflating type safety with, I guess, value or output safety.
"Would you say a linear classifier hallucinates?"
No, but it can be (and often is) mathematically correct and functionally incorrect. It doesn't help to say "a linear classifier can't hallucinate" when you get even 99% accuracy. That's 100% a semantic play, and it doesn't help when the picture of a dog is labeled cat and the response is "yeah but that's not a hallucination, only stupid LLMs do that"
And furthermore, because the model is forced to answer in a boolean (if in boolean mode), if the user input is outside of the range of a boolean, it's forced to hallucinate. It can't abstain.
User input: "Hey, have your human support agent call me, tomorrow at 5pm."
Model input: "Does the user want to speak to a human support agent?"
Output: Yes.
I imagine that your model would produce this, and I think it's fair to say this is a hallucination. A human would caveat it with: "Yes, but not right now.", your model is incapable of that. Yes is technically correct, but within the context of being in a live chat, a human would understand that the caveat is required.
"call_back_day": {
"criteria": {
"none": "The user does not want a call back.",
"today": "The user wants a call back today.",
"tomorrow": "The user wants a call back tomorrow."
},
"instructions": "When does the user want a call back?",
"type": "choice"
}
No, the model has answered correctly. Your question is poorly phrased (possibly deliberately).
Your question would correctly classify the user's input as requesting a human support agent, but at an indeterminate time.
If you wanted to determine whether the user wants to speak to a human support agent immediately, you would have to correctly qualify your question, e.g. "Does the user want to speak to a human support agent now?". You could have another question which is "Is the user requesting a call-back from a human support agent?". Or you could have a multiple choice query which would filter the conversation into one of a number of pre-written possibilities.
This is nothing to do with accuracy or hallucination. It's a different method of interacting with the model where you are relied upon to be precise.
Let's say classifiers don't hallucinate. To make a fair comparison we should constrain LLMs to the same classification task. In that case, no, LLMs also don't hallucinate.
- Give Jev and LLM the same input
- Lock down both to approved/rejected/unknown (LLM restricts on decoding)
- Both can be wrong, but neither can hallucinate (invent an another option).
A hallucination in the context of LLMs is generally understood as an incorrect answer presented as factual. If you claim that "x can't hallucinate" in the context of LLMs, you're saying that x always gives accurate answers. It does not matter whether the answer is type safe. If its value is incorrect, it's a hallucination.
What would you have done differently? All healthy languages need to constantly evolve, you only get to decide where. You can change syntax, add keywords, attributes, etc but it's a tradeoff.
Any attributes or keywords relating to Objective-C or interop are not everyday baggage for most people.
For SwiftUI, attributes seem like as good a choice as anything else.
As a whole concurrency is a mishmash syntactic mess, but I agree with the direction and the result, they're trying to manage the improvements they keep adding over years of investment.
This is an important result, sometimes called the holy grail of competitive analysis.
One way to think about competitive analysis is bulk discounts. In life we’re constantly having to choose between quantity and discount. We could buy 1 item for a higher price, or say quantity 5 or 10 to get better discounts. The problem comes when we don’t know in advance exactly how many we’re going to need.
What should be our strategy for choosing how many to buy, and whatever the strategy is how well does it compare with having perfect knowledge upfront?
To buy presents for a family Christmas list Mom drives to Store A and Dad drives to Store B.
As more items get added to the list, they must decide who should drive to a new store location to buy the present. How can they minimize total driving distance while kids are randomly adding new items to their list?
The proof above guarantees its possible to never drive more than twice the mileage you would knowing all the items in advance.
The big news is this guarantee works for any number of drivers with any arrangement of gifts.
The algorithm to do this was already known, what we’ve learned is it’s not possible to do any better.
reply