Nowadays we use LLMs mostly for doing agentic-based work. LLMs new Pareto frontier only make the headlines if they push the boundaries on benchmarks that are deterministic tasks. So models are encouraged to focus on these deterministic tasks that are, in nature, structured texts. I think that this makes models more “plastic” or “polished”, as opposed to natural and pleasant to read. User-based benchmarks, like LLM Arena, are for me the best we can do in order to rank models in this way, but come with its own drawback (subjective evaluation, prone to spam or techniques to promote a giving model).
For those that don't know about this. Phi was announced with a paper called "Textbooks are all you need". What they did was use GPT 3.5 and created synthetic textbook chapters and exercises.
They also did some more interesting work like showing very small models can be coherent as long as you have very simple children's book style training data (TinyStories is pretty famous).
Lots of these ideas are still used. Learning facts at scale with active reading is an ICLR 2026 paper from Meta AI that does a lot of similar work.
Not to demerit the recording, but I felt more nostalgic for the last sentence of the article "Sometimes, the internet is good" than for the musics itself.
We all know it... but I think they were very bold in this warning about using your private messages to train public models.
_Your messages with AIs will be used to improve AI at Meta. Don't share information, including sensitive topics, about others or yourself that you don't want the AI to retain and use_
"A watchdog kernel thread monitors RAM and NVMe pressure and signals userspace before things get dangerous." - which kind of danger this type of solution can have?
reply