AI-Native Academic (1): You've Got to Build Your Own Tools with Frontier AI
How we used to learn
Think of how you learned before ChatGPT. First you had to select what to learn, because human attention is scarce and you cannot read and screen everything. So we decided what was worth learning by trusting authorities and intermediaries: curated resources such as textbooks and courses; the professor at the front of the class or on the screen; rankings and third-party evaluations, the “best books” lists; the journal lists that told us which venues counted; the open-source community, where you posted your question on Stack Overflow or the Stata forum and waited for someone who knew;1 and so on. Somebody else had read widely, made the evaluations and the judgment calls, and we learned from their wisdom.
Second, learning was bottom-up. You learned the fundamentals before you learned what sat on top of them. To learn machine learning you learned linear algebra first, and how to print “hello world” in Python, and then a great deal more before you could use a Python or R package to run a topic model over ten thousand annual reports, or classify the sentiment of a million reviews.
Third, we adjusted ourselves to our tools. Stata, R, MATLAB, LaTeX: each was a finished instrument with a syntax of its own, and we learned the syntax before the instrument would do anything for us.
What has changed
There was nothing wrong with any of this. It is just that, in September 2026, the reality has changed.
How we learn
First, how we learn has to be re-engineered for this reality: where it was bottom-up, we now go top-down. Gabriel Petersson dropped out of high school in 2019 and was an OpenAI research scientist by 22. He learned machine learning from a project down, asking ChatGPT about each concept as he hit it. “You start with a problem, you recursively go down,” he says, and then “it doesn’t need to go bottom up anymore.”2 We used to have to learn the fundamentals of a piece of software, the syntax of Stata or R, before we could do anything meaningful with it; now we have to know what we want to do with it, and AI helps us learn the rest, top-down. Frontier models can do Stata and R, LaTeX and MATLAB, and the programs themselves have let AI in: MATLAB has had a Copilot since October 2025, RStudio’s daily build got Posit’s assistant in March, Overleaf’s AI writes the LaTeX for a table from a text prompt, and for Stata the bridges to the models have come mainly from its users, in the community’s long tradition of contributed tools.3 They read the manual and write the code at a speed and accuracy a person can hardly compete with.
The head of Claude Code said in June that he had not written a line of code by hand in about eight months, and that one developer had rewritten the Bun runtime from one language to another in six days, work he estimated would once have taken about a year.4 At Anthropic itself, as of May, more than 80% of the code merged into the company’s codebase was written by Claude, by the company’s own count.5 With GPT-6 Astra this goes further: software that used to be hard to learn, the model can now just operate. Take Blender: OpenAI’s launch demonstrations show Astra modeling a house in it and carrying the scene into a walkable Unreal Engine environment, and OpenAI says Astra “marks a new frontier in the speed, accuracy, and safety of computer use.”6
What a tool is
Second, the meaning of “tool” has changed with frontier AI. Now that AI can do complicated tasks on its own, the AI itself is a tool, and an AI agent is also a tool for our research work. A tool, in this piece, is anything that gets a piece of your research work done. The model is one: give it the task, and it does it. An agent built on the model is another: a model, a set of instructions, and access to the things it needs, packaged so that it carries out the task without you at the keyboard. Workflow, skill, agent: the names vary, and each is a way of getting work done that you write down, run, share and change. Anthropic’s engineers class workflows and agents alike as “agentic systems,” calling workflows “systems where LLMs and tools are orchestrated through predefined code paths” and agents “systems where LLMs dynamically direct their own processes and tool usage”;7 the open Agent Skills standard, started by Anthropic and adopted by a growing number of agent products, packages “procedural knowledge” into folders that agents “load on demand,” to “turn multi-step tasks into consistent, auditable procedures”;8 and Andrej Karpathy, who calls the shift Software 3.0, says that “your prompts are now programs that program the LLM.”9
People have been building such tools. A professor at Arizona State gave Fable 5 in Claude Code a research question and the materials from an earlier one-shot attempt, and let it run: by the model’s own audit, more than 80 agents, about 2,000 tool uses, more than 120 primary sources and nearly fifteen hours later it returned a dense 53-page report. After two rewrites for readability, he judged the resulting paper far better than the one-shot answer, and said he would need days to assess it.10 A neuroscientist at the Allen Institute built a multi-agent review template with about twenty custom skills and now has about ten reviews, many of more than a hundred pages, with citations checked by reviewer agents; Anthropic says such a review could take Lecoq’s team as many as two years.11 And in August a team of 93 authors led from Harvard and MIT released MatrAIx, an evaluation system built on 8.3 billion persona records, some synthetic and some derived from human-authored profiles, on a shared schema of 1,290 attributes; large language models animate them as agents that fill in surveys, talk to chatbots, browse websites and use apps. They report 18,189 evaluation trials across eight tasks from a library of 1,010, and a subset of about one million personas is public for non-commercial research.12 So, with frontier AI, everyone can build their own tools.13
David Rogier, who runs MasterClass and builds his own working tools, put it this way in June: “you’re able to create the workflows and tools, anything that you now want. So now it’s about what’s the motivation to do it, what’s the skills to go do it.”14 Hence, the building is the easier half, because Fable 5.1 or GPT-6 Astra can easily do it (oftentimes now the issue becomes over-building or over-engineering). The harder half is to identify the need, a gap in your skills, in your workflow, in what your tools can do, and then to use AI to build the tool that covers it. It takes more agency here: to reflect on one’s own workflow, on how things could get done differently with frontier AI capabilities, on which gap to fill first, and on how to evaluate and make sure the tool does the job.
Keeping up with the models
Third, the AI models keep evolving, and we have to keep learning. OpenAI shipped GPT-5.4 in March, GPT-5.6 Sol in July and GPT-6 Astra in September; Anthropic shipped Claude Fable 5 in June and Fable 5.1 in September; and OpenAI keeps retiring older models from ChatGPT so that it can better serve the newer ones.15 Each jump changes what the model can do on its own, and so what we should still tell it. When Fable 5 came out, Anthropic’s own prompting guide said that skills written for the earlier models “are often too prescriptive for Claude Fable 5 and can degrade output quality,” and that a capability jump of this size is “also a good prompt to re-evaluate which instructions, tools, and guardrails are still needed.” Boris Cherny, who built Claude Code, said the same of the product itself: “Every time that a new model comes out, we delete a bunch of the system prompt,” and of Opus 5 he said that much of it “was correcting for these behaviors that the model should have known, but it didn’t. Now Opus 5 just does it. So yeah, we deleted 80% of the system prompt.”16
The way we use AI changes as well: OpenAI trains GPT-6 Astra to divide and delegate work to subagents, and says it is generally better than earlier models at staying coherent during long tasks.17 So for us the implication is this: each time the models make a breakthrough, we have to rethink how we organize our work around the new capabilities. A tool built on one model, with prompts and workflows tuned to its habits, may need an overhaul when the next model comes, and the next one comes within months. A list of “the ten must-have academic tools” built on GPT-5.5 or Opus 4.8 is most likely stale by now unless someone has updated and overhauled it. Keeping up has to be done often, or you are left behind.
Human agency in the AI era
So, facing all these new realities in September 2026, we need more agency as learners and researchers. How to use a frontier model to become a better learner than we could be before is itself something we have to learn. It is a skill, and like any skill it is learned by doing. As Fei-Fei Li, who built ImageNet and now runs World Labs, said in a recent interview when asked what people should prepare for in the workplace ten years from now: “Agency. I think AI will give people more agency.” About seven minutes later she put it as an instruction: “in the face of a technology that is so cognitively advanced, be brave, have your human agency, and command that technology.”14
The AI Academic Hub












So what I want to do here is to curate a knowledge vault for our agents. Anyone can now build their own tool, or take an existing one and bend it to their need, with a frontier model doing the building. But one thing is still needed: knowing what is out there to build from, and knowing it before it is stale. The AI Academic Hub is my answer to that, for myself, and for my fellow researchers who want to adapt to an AI-native way of building tools for their research.
What it is
A knowledge vault for research work, in the pattern of the LLM wiki that Andrej Karpathy set out in April: an AI builds the pages from sources, maintains them, and carries what it learned from one question into the next. In his words, the model “incrementally builds and maintains a persistent wiki” while “You’re in charge of sourcing, exploration, and asking the right questions”. The result is “a persistent, compounding artifact” that “keeps getting richer with every source you add and every question you ask”.18 The searches are re-run as the field moves, and every page carries the date it was last checked.
It is a wide survey and database of repositories, articles and datasets, built on searches that serve the skills, tasks and workflows academics do often, such as literature review, theorization, data analysis, and reporting and communicating research outputs, and each work is placed under the task it serves. For each task it shows what exists and what each work does. The searches reach into open source, working papers and repositories, where the state of the art now often appears first; in AI especially, conference papers can lag.
For you and your AI, I used a rigorous, social-science-backed search method and a multi-agent workflow to collect the building blocks, for you to adapt to your own needs and build your own tool. It carries no ratings and no human picks, because every need can be unique, and a list calibrated to nobody’s voice flattens everything toward the average; the customization is yours, and you can do it with your AI.
How I built it
I built a search workflow, powered by social-science methods and run by many models and many agents, and a record that holds what it finds, so that the output can be checked and the search can be replicated. An AI chat answer is one draw from a distribution: ask twice and you may get two lists, and either may carry a fabrication.19 A workflow reaches wide, writes down what it looked at, and can be repeated and checked by another person or another model.
First, the search follows social-science canons. Each campaign is run like a systematic review: a written scope fixes the question; a search log records every query; a screening sheet records every lead and the reason it was kept or dropped; an evidence ledger holds what was retained, each item with an exact locator; a findings file states what was found and what was not, with a plain label of how far the search reached; and the bundle is sealed and dated.
Second, it runs as a workflow. Frontier models can run many agents at once and stay on long tasks: Claude Code coordinates dozens to hundreds of agents in one scripted workflow, and GPT-6 Astra is trained to divide work among subagents and is generally better than earlier models at staying coherent during long tasks.17 The agents follow the workflow I wrote, step by step, and they keep going. This is a scale of work beyond any one person.
Third, it uses more than one frontier model, because the models complement each other. Two model families given the same search question overlapped on only about a fifth of everything they found between them; the union was the wider map.20 A study this year of LLM ensembles argues that the value of a model “lies in its complementarity with others.”21 So in my runs a second family extended the coverage, and a second model that did not do the search checks the result. That check, with the fixed question, the recorded screening and the dated pages, is how I engineer against hallucination and fabrication. Every retained finding links to its source, and a person approves what is published.
How to use it
You get a larger set of traced options, each with what it is, where it lives, its license, and what it takes to use. Those are components; you do the composing, with the help of your AI. Adapt them to your own workflow and habits, and choose what fits your work. That last step is what the title means. The hub does the first half of the problem, what is out there and whether it still holds. The second half was never anyone else’s to do, because only you know what your work needs. Take the components that fit, and with your own AI, build the tool that fits you. The building is the easy half now. You’ve got to build your own tools with AI, and you can.
Sources
-
R. Maria del Rio-Chanona, Nadzeya Laurentsyeva and Johannes Wachs, “Large language models reduce public knowledge sharing on online Q&A platforms,” PNAS Nexus 3(9), pgae400 (2024). Within six months of ChatGPT’s release, posting activity on Stack Overflow fell by about 25% relative to the platform’s Russian- and Chinese-language counterparts and to mathematics Q&A sites, with no significant change in post quality, measured by peer feedback (votes). The study covers Stack Overflow; the Stata forum is my own example. The example marks both the before and the after: the forum was where one went, and then posting there fell. ↩
-
Sigil Wen, “High School Dropout to OpenAI Researcher - Gabriel Petersson Interview (Extraordinary)”, Extraordinary podcast on YouTube (November 2025); the quotations are at 12:04 and 12:40 of the recording, checked against YouTube’s automatic captions. The 2019 dropout and his title are from Lee Chong Ming, “A high school dropout who got hired at OpenAI says he used ChatGPT to learn Ph.D.-level AI,” Business Insider (28 November 2025), reporting the episode; his age and nationality are from Preston Fore, Fortune (29 March 2026). ↩
-
MathWorks, “MathWorks Launches Generative AI-powered MATLAB Copilot to Boost Productivity and Accelerate Development for Engineers, Scientists, and Researchers” (7 October 2025): “a generative AI assistant for MATLAB,” available in Release 2025b. Joe Cheng, Nick Rohrbaugh and Sara Altman, “Introducing AI in RStudio” (Posit, 5 March 2026): Posit Assistant is “a conversational agent that lives in your RStudio workspace” that can “visualize patterns, reshape datasets, debug errors, build Shiny apps, explain code,” available that day in the daily build of RStudio through a paid Posit AI subscription; “Using RStudio without AI will always remain free.” Overleaf, AI features, read 17 September 2026: the table and equation generator “Generates LaTeX tables and equations from images or simple text prompts,” and TeXGPT “Generates figures and other LaTeX code”; the page adds, “It won’t write your paper for you, but TeXGPT can help you create it faster.” For Stata the bridges have come mainly from the community, which the page places in its “long history of extending the software’s capabilities through community-contributed tools”: StataCorp’s own newsletter, “Community corner: AI tools for Stata” (Stata News, volume 41, number 2, 2026, read 17 September 2026), lists a community-contributed chatgpt command, two user-written projects (Song Tan’s Stata-MCP, which lets AI agents run regressions and analyze data, and Thomas Monk’s mcp-stata server), and a tutorial on using Claude Code with Stata, and calls them “community-led projects.” It also links a Stata Blog post by StataCorp’s Director of Statistical Outreach, Chuck Huber, on writing Stata commands that call ChatGPT, Claude, Gemini and Grok (7 October 2025). All are vendor pages. ↩
-
Boris Cherny in conversation with Jeremy Kahn, “From Prompts to Power”, Fortune Brainstorm Tech, Aspen (8 June 2026), video on Fortune’s site; quotations checked against Fortune’s caption file. At 4:33 he says, “I haven’t written a line of code by hand in, you know, I think 8 months now.” At 19:02, on Jarred Sumner’s rewrite of Bun from Zig to Rust with Opus 4.8 and dynamic workflows, he says, “This would have taken probably a year of work before if he did it in 6 days.” Cherny’s own claims. In the same talk he cited Anthropic’s figure of about 8x more code; Anthropic’s own report calls its 8× lines-of-code figure “almost certainly an overstatement of the true productivity gain.” Fortune’s report of the talk is Nick Lichtenberg, “The head of Claude Code hasn’t ‘written a line of code by hand’ in 8 months,” (11 June 2026). ↩
-
Marina Favaro and Jack Clark, “When AI builds itself”, Anthropic Institute (undated; first archived 4 June 2026): “As of May 2026, more than 80% of the code we merge into Anthropic’s codebase was authored by Claude.” The company’s own figure, in a report arguing for the option to slow frontier development. ↩
-
OpenAI, “GPT-6 Astra: A new generation of intelligence” (3 September 2026): “GPT‑6 Astra marks a new frontier in the speed, accuracy, and safety of computer use”; with its demonstrations labeled “Blender model” and “Unreal Engine walkthrough,” the page says “GPT‑6 Astra models a house in Blender and turns it into a walkable scene in Unreal Engine 5.” A vendor page. ↩
-
Erik Schluntz and Barry Zhang, “Building effective agents”, Anthropic (19 December 2024): “At Anthropic, we categorize all these variations as agentic systems, but draw an important architectural distinction between workflows and agents”; “Workflows are systems where LLMs and tools are orchestrated through predefined code paths. Agents, on the other hand, are systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks.” A 2024 text, kept for its definitions; the byline now reads “Erik S.” ↩
-
Agent Skills, the standard’s own site, read 17 September 2026: “Agent Skills are a lightweight, open format for extending AI agent capabilities with specialized knowledge and workflows”; “a skill is a folder containing a SKILL.md file” with “instructions that tell an agent how to perform a specific task”; skills package “procedural knowledge” into “portable, version-controlled folders that agents load on demand”; one listed benefit reads “Turn multi-step tasks into consistent, auditable procedures”; the format was “originally developed by Anthropic, released as an open standard, and has been adopted by a growing number of agent products,” among them Claude Code, ChatGPT and Codex, Gemini CLI, Cursor, GitHub Copilot and VS Code, as the site’s Client Showcase lists them. ↩
-
Andrej Karpathy, “Software Is Changing (Again),” keynote at Y Combinator’s AI Startup School, San Francisco (17 June 2025). In YouTube’s automatic captions, at about 3:19, he says the new paradigm is “worth giving it a new designation of software 3.0” and that “your prompts are now programs that program the LLM.” A 2025 talk, kept as a definition. ↩
-
Andrew Maynard, Professor of Advanced Technology Transitions at Arizona State University’s Thunderbird School of Global Management (About page), “A quick update on using Claude Fable 5 for research,” The Future of Being Human (12 June 2026), his own account of his own run, self-reported. He gave Fable 5 in Claude Code the original prompt and one-shot paper, reviews of it, his own feedback and a folder of cited documents. According to “Fable’s own audit,” the project used over 80 agents, “called on the use of 2000 tools,” verified over 120 primary sources, downloading and auditing them, ran for nearly 15 hours and consumed over 8 million tokens, nearly half of them on verification; it produced a dense 53-page report on Mythos-class models and higher education. Of the stand-alone paper Fable wrote from it after two further rewrites for readability, he says: “Even on a quick read, this paper is substantially better than the one produced from the one-shot prompt. By a long way.” He adds that he “would need to spend days with this — probably more” before he was sure of its value, and that the earlier one-shot paper tended to rely on secondary sources. ↩
-
Anthropic, “Claude Science, an AI workbench for scientists, is now available” (30 June 2026): Jérôme Lecoq, a neuroscientist at the Allen Institute, used Claude Science to build a multi-agent “computational review template” with about 20 custom skills; the page says he “now has about 10 reviews, many more than 100 pages, with citations that were checked over by reviewer agents,” and that before Claude Science “it could take Lecoq’s team as many as two years to write such a review.” A vendor page. ↩
-
Xiaomin Li and 92 co-authors, “MatrAIx: Simulating the World with 8.3 Billion Persona Agents,” arXiv 2608.04205 (4 August 2026); the two organizers, Xiaomin Li and Yuexing Hao, are at Harvard University and MIT (appendix A). Code at github.com/MatrAIx-ai/MatrAIx-Persona-8B, MIT license; the Persona 1M coreset on Hugging Face, released for non-commercial research use only, and the dataset card says the repository’s MIT license covers only the software in that repository. The abstract reports 8.3 billion persona records on 1,290 categorical dimensions, either sampled from a dependency graph or derived from human-authored profiles; four environments (Survey, AI Chatbot, Web, App); a library of 1,010 tasks in more than 25 domains; 18,189 trials across eight representative tasks, with persona agents powered by Claude Opus 4.8, GPT-5.5 and Claude Haiku 4.5; and, in a 400-trial controlled study, the declared behavior expressed or correctly suppressed in 366 trials (91.5%), with Claude Opus 4.8 as the acting agent (appendix I). The public coreset holds 999,847 personas, 599,847 human-grounded and 400,000 synthetic. The numbers are the authors’ own. ↩
-
A hobbyist who had wanted to design a circuit board for years described one to Claude Fable 5 in plain English, let it do the schematic, the parts and the routing, sent the files to a factory, and got back five assembled boards for 130 euros, displays extra; the first board plugged in worked. A reader of the same post had Sol design a mini Bluetooth keyboard meant to clip onto a phone. One developer had Fable 5 build a 4×4-kilometer procedural world that runs in the browser, about 21,000 lines of code that the repository says are roughly 99% the model’s; another shipped the first version of a browser MMO with nine classes and a five-player dungeon in two days, on the side. Sources: A6M-Zero, “This PCB is brought to you by Fable 5” (4 September 2026): an RP2350A-based board with a 1.54-inch e-ink display, four buttons and expansion pins, a four-layer PCB in KiCad, under two rules the author set, “No manual edits or verification of the board” and that every problem before manufacturing would be solved by Fable; 65 initial design-rule errors, two component-package mismatches caught only after upload, 49 connections left by the Freerouting tool and then routed by hand by Claude; five fully assembled boards for €130, with the e-ink displays bought separately; the first board plugged in “was recognized and it was ready to be used.” The keyboard is from the Hacker News thread on that post (14 September 2026), commenter sottol, whose boards were “done with Sol”: a mini Bluetooth keyboard with Kailh PG1316S switches, meant to clip onto a Pixel phone once a housing is designed and made, about $70 to produce plus $80 for shipping, taxes and fees, generated as Python that writes KiCad files; fabricated and awaiting shipment at the time; on 17 September the commenter reported “I just got my PCBs,” with the switches and microcontroller still to be soldered. Both are self-reported by their builders. Braffolk, fable5-world-demo (GitHub, MIT license, 722 stars when read on 17 September 2026): a procedural 4×4 km open world in WebGPU, about 21,000 lines of TypeScript, “built roughly 99% by the model, with minimal human steering”; the human partially wrote one document of targets and constraints. ls-sadboy, comments in the Hacker News thread “Mmorpg World of ClaudeCraft, vibe coded with Fable 5” (12 June 2026), the builder’s own words: “a vanilla-WoW-flavoured micro-MMO in the browser: nine classic classes, three zones, a 5-player instanced dungeon,” built “over a couple of days, on the side” using “like 93% of my Max plan over the course of 2 days to get the initial version running,” released under MIT. Both are self-reported. ↩
-
Marina Mogilko, “Godmother of AI: In 10 Years There Will Be Only 2 Kinds of Workers,” Silicon Valley Girl Podcast, with Fei-Fei Li of World Labs and David Rogier of MasterClass (19 June 2026), transcript on the episode page. The page’s own description says Li “built ImageNet” and “now runs World Labs,” and calls Rogier “the CEO of Masterclass,” “known for building his own AI-powered productivity tools rather than relying on commercial software.” In the YouTube recording, Li’s “Agency” answer starts at 21:46 and the “be brave” line comes at 28:42. Asked to picture the workplace of ten years from now and what people should prepare for, Li answers: “Agency. I think AI will give people more agency. A lot of the future of work I think will rely on people who know how to use these tools in a very effective way.” The “be brave” line follows in the exchange on specialists and what Rogier calls the “high-agency generalist.” Rogier’s line comes after he describes the apps he has built for himself, among them a to-do list that will not keep a task for more than a day and a half: “the cost of making an app dropped from months to like a weekend.” Two practitioners in conversation; the show’s own note reads the divide as one “between high-agency workers who embrace AI and those hesitant to engage with it.” ↩ ↩2
-
OpenAI Help Center, “Model Release Notes”, read 16 September 2026: GPT-5.4 Thinking (5 March 2026), GPT-5.4 mini (18 March), a GPT-5.5 Instant update (28 May), GPT-5.6 Sol (9 July); the same page records that GPT-5.1 models were removed from ChatGPT on 11 March 2026 and that OpenAI o3 would be retired from ChatGPT on 26 August 2026 after a 90-day sunset: “These changes apply to ChatGPT only; there are no changes to the API.” The GPT-6 Astra system card is dated 3 September 2026. Anthropic’s Claude Fable page, read the same day, gives 9 June 2026 for Claude Fable 5 and 1 September 2026 for Claude Fable 5.1. ↩
-
Anthropic, “Prompting Claude Fable 5”, Claude Platform documentation, read 17 September 2026: “Capability improvements at this level are also a good prompt to re-evaluate which instructions, tools, and guardrails are still needed,” and, under Recommended scaffolding changes, “Skills developed for prior models are often too prescriptive for Claude Fable 5 and can degrade output quality. Review and consider removing older instructions if default performance is better.” Boris Cherny, interviewed by Diana Hu at Y Combinator’s Startup School 2026, “Boris Cherny: We Cut 80% of Claude Code’s Prompt” (Y Combinator, 27 July 2026; in the YC Library as “Boris Cherny: Building Claude Code”), at 3:51: “Every time that a new model comes out, we delete a bunch of the system prompt, change a bunch of the system prompt, we change the set of tools all the time, we change the prompts for the tools all the time,” and at 4:27: “Now Opus 5 just does it. So yeah, we deleted 80% of the system prompt.” At 6:58 he adds, for people who use Claude Code: “every 6 months delete your Claude MD. Delete your skills. Delete your hooks. See what the model does and it might surprise you.” Fillers are omitted from the quotations. Anthropic’s own account of the cut is Thariq Shihipar’s “The new rules of context engineering for Claude 5 generation models” (Claude blog, 24 July 2026): “We removed over 80% of Claude Code’s system prompt for models like Claude Opus 5 and Claude Fable 5 with no measurable loss on our coding evaluations.” Cherny’s and Shihipar’s figures are self-reported. ↩
-
Anthropic, “Orchestrate subagents at scale with dynamic workflows,” Claude Code documentation, read 16 September 2026: a dynamic workflow is “a JavaScript script that orchestrates many subagents at once,” and the page gives its scale as “Dozens to hundreds of agents per run.” OpenAI, “Model guidance,” API documentation, section “Using GPT-6 Astra,” read the same day: “GPT-6 Astra is trained to be able to divide and delegate work to subagents that work in parallel,” and it “is generally better than GPT-5.6 Sol and earlier models at staying coherent during long tasks.” Both are the vendors’ own descriptions of their products. ↩ ↩2
-
Andrej Karpathy, llm-wiki.md, a GitHub gist created 4 April 2026: an LLM “incrementally builds and maintains a persistent wiki” of pages, an index and an append-only log; “the wiki is a persistent, compounding artifact” that “keeps getting richer with every source you add and every question you ask”; the LLM writes and maintains it, and “You’re in charge of sourcing, exploration, and asking the right questions.” The gist’s own checks are a periodic lint pass, in which the LLM is asked to “health-check the wiki,” and answers given “with citations.” The hub follows that pattern and adds a check by a second model that did not do the search, which the gist does not describe. ↩
-
The labs say so themselves. OpenAI’s system card for GPT-6 Astra reports fewer factual errors than GPT-5.6 Sol and still measures hallucinations, on a test built from conversations that users had flagged as containing factual errors. An independent benchmarking firm found that Claude Fable 5.1 gave the most accurate answers it has measured on its knowledge test, and also that when the model did not know an answer, it gave a wrong one nearly three times in four. Each answer is also one draw from a distribution: Anthropic’s glossary says “identical inputs may produce different outputs across API calls,” and OpenAI’s API guide says model outputs “may differ from request to request.” Sources: OpenAI, “GPT-6 Astra System Card” (3 September 2026), section 7, Hallucinations. Astra “makes substantially fewer factual errors than GPT-5.6 Sol.” The evaluation uses de-identified ChatGPT conversations that users of prior models had flagged as containing factual errors, and OpenAI cautions that these are especially hallucination-prone cases whose error rates “should not be interpreted as hallucination rates observed in production.” Artificial Analysis, “Claude Fable 5.1 tops the Artificial Analysis Intelligence Index” (1 September 2026), an independent benchmarking firm that also ran pre-release evaluations for Anthropic. On its AA-Omniscience test, Fable 5.1 at maximum effort attempted 93.4% of questions and recorded the highest accuracy the firm has measured, 67.2%; of the questions it did not answer fully correctly, it gave an incorrect answer 72.6% of the time, against 63.6% for Claude Fable 5; the rest were partial answers or abstentions. The firm defines this hallucination rate as “the proportion of incorrect answers out of all non-correct responses,” partial answers and abstentions included (evaluation page, read 19 September 2026; benchmark note, 16 November 2025). Anthropic’s own system card for Fable 5.1 and Mythos 5.1 (1 September 2026) carries the lab’s own honesty evaluations (section 6.5, run on the Mythos 5.1 snapshot of the same model); I cite the independent figures here. Anthropic, Glossary, Claude Platform documentation, and OpenAI, Advanced usage, API documentation, both read 19 September 2026. ↩
-
My own record: the comparison of the two HUB-P1 bundles, one produced by a Codex seat and one by a Claude seat given the same scope on 15 September 2026, filed as pull request 5 in my scout repository. Of the retained identities, 39 were found by both, 36 only by the Codex seat and 101 only by the Claude seat, a Jaccard overlap of 0.22. Two campaigns, one per model family, that also differed in tools, parallel lanes and time (about 35 against 90 minutes); the counts do not separate the model from the rest. ↩
-
Yichi Zhang and five co-authors, “Mixture of Complementary Agents for Robust LLM Ensemble,” arXiv 2605.24048 (21 May 2026): the abstract treats the choice of which models to feed a summarizing model as a selection problem “where the value of an LLM lies in its complementarity with others,” tests greedy algorithms that “assess complementarity using a small labeled set,” and reports methods with the best performance-cost trade-offs. Its experiments use eight models cited to 2024 and 2025 reports, on three reasoning benchmarks. ↩
© 2026 Joy (Zhao) Zheng. All rights reserved.