Digestion Is Not Absorption: Karpathy's Output Ladder, Feynman's Layer, and an Agent Loop That Closes the Gap
What Karpathy's 2026-10-02 output-format post optimizes (rendering knowledge for one reader) and what it leaves out (any check that the reader actually absorbed it). The Feynman technique supplies that missing layer from the opposite direction, and an LLM agent can run both as one loop: scope a task-sized knowledge subset, render it, test the reader on exactly that subset, re-scope from the results. Includes a concrete playbook with per-phase prompts and sign-off criteria.
A widely shared post by Andrej Karpathy (x.com/karpathy/status/2105819303471976479, 2026-10-02) opens with: “We’ll be spending a lot more time trying to understand the outputs of language models.” It proposes output formats for that understanding, ranks them, and closes with the claim that as intelligence and code become abundant, “you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before.”
This post asks the next question the post does not answer: how do you know anything was actually digested? The short answer: the post has no mechanism for that — and a different, older technique does. The two compose into one agent loop.
Method note: the tweet text quoted below was retrieved via the api.fxtwitter.com mirror (x.com itself is login-walled), and cross-checked against Kruno Golubić’s writeup. Claims are labeled as sourced fact, interpretation, or design proposal.
1. What the Karpathy post optimizes: rendering
The post is a ladder of output formats, each presented as “even better” than the last:
- Writing — ask the LLM to explain something in ASD-STE100, “a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable.” Because the spec is stringent, he softens it: “ask for ‘80% of the way to ASD-STE100’”.
- Diagrams / images — “easier to process, parse, and understand.”
- Interactive HTML pages — “a beautiful, interactive webpage … beautiful experiences, animations.”
- Explainer videos — “the output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic.”
The unit of improvement is the artifact: how well the material is rendered for this reader, this question, this moment — generated once, consumed, thrown away, because generation is now cheap.
One factual note on originality: the ASD-STE100 prompt trick predates the post. Andrew Carr tweeted “The fix for this is to say: only report to me in ASD-STE100 Simplified Technical English” on 2026-07-27 (x.com/andrew_n_carr/status/2081534245370314816, verified via the same mirror), and per Golubić’s account the instruction spread through prompt-sharing and several community “skills” in the following weeks. What the Karpathy post added was reach — and the format ladder around the trick.
2. Counter-argument: easier to digest does not grant better absorption
Every quality claim in the post is about the artifact — “easier to process, parse, and understand”. None of them touch the reader’s state. A re-scan of the full verified text finds zero occurrences of test, quiz, evaluate, verify, assess, recall; the only near-hit is “oversight” (once, in the summary bullet about work rising “into oversight and understanding”). Digestion is assumed, never measured.
The learning-science evidence says the reader’s side is governed by something rendering cannot supply:
- Testing beats re-exposure. Roediger & Karpicke (2006), “Test-Enhanced Learning”, Psychological Science (doi:10.1111/j.1467-9280.2006.01693.x): taking memory tests improves long-term retention versus restudying. Re-reading and re-watching — exactly what a polished bespoke artifact invites — is the weak intervention; retrieval is the strong one. (Paper identity verified via Crossref.)
- Practice testing is the top-rated technique. Dunlosky, Rawson, Marsh, Nathan & Willingham (2013), “Improving Students’ Learning With Effective Learning Techniques”, Psychological Science in the Public Interest (doi:10.1177/1529100612453266) is the standard review ranking technique effectiveness; its top-utility rating of practice testing is widely cited. (Precision note: the paper’s identity is verified via Crossref; the abstract text could not be fetched this session — PubMed and the publisher both bot-walled — so the exact rating wording is not quoted here.)
- Interaction, not rendering, drove the AI-tutor gains. Kestin, Miller, Klales, Milbourne & Ponti, “AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting”, Scientific Reports 15, article 17458 (nature.com/articles/s41598-025-97652-6): “students learn significantly more in less time when using the AI tutor, compared with the in-class active learning” — and the tutor’s design was built on active-learning pedagogy (questions and feedback), not on prettier pages.
- The fluency illusion (interpretation, anchored to the above): smooth, personalized, well-rendered consumption generates familiarity, and familiarity feels like understanding. A bespoke explainer maximizes fluency — precisely the condition under which the illusion is strongest. The better the ladder works, the more silently it can fail.
So: the ladder raises the ceiling of exposure quality; absorption is untouched, and the binding constraint moves to verification.
3. The old way: Feynman serves the absorption layer
The Feynman technique, as popularized by fs.blog, has four steps: select a concept and map your knowledge → teach it to a 12-year-old → review and refine → test and archive. Its premise: “complexity and jargon often mask a lack of understanding.”
Why it targets exactly the layer the output ladder leaves open:
| Karpathy’s ladder | Feynman | |
|---|---|---|
| Who generates the explanation | the LLM | the learner |
| What it measures | nothing (assumes digestion) | the learner’s actual state |
| Layer | rendering (input quality) | absorption (output evidence) |
The learner must generate the explanation. Generation is retrieval practice in disguise (the testing effect, §2), and the points where the explanation breaks down are the coordinates of what is not yet understood — a diagnostic instrument, not a study habit. It is also the classic countermeasure to the fluency illusion, applied from the opposite direction.
4. Playing Feynman on an LLM/agent setup
(Design synthesis from here on; the orchestration is a proposal, the per-element evidence anchors are cited above.)
The user who asks about “calculus” almost always needs a task-scoped subset — the differentiation machinery required by one specific problem — not the domain. The agent’s first job is to find that subset; its second is to verify it landed.
Stage A — interactive scoping: what does the user actually want, and how much do they care
Before teaching anything, the agent interviews and writes down a study contract:
- Subset — the minimal knowledge actually required by the triggering task.
- Not-in-scope list — explicit, named exclusions. This is what keeps “cheap generation” from silently becoming over-generation.
- Success criterion — what the user will be able to do when done (a behavior, not a vibe).
- Depth budget — how much the user actually cares: a ten-minute pragmatic brief versus a full mental model.
The interview is the mechanism that converts abundance into relevance. The contract is confirmed by the user before Stage B starts; if the user cannot state what they care about, the agent offers 2–3 presets instead of guessing.
Stage B — the dive/test loop: render, ask, diagnose, correct, re-scope
The agent alternates rendering and retrieval, and every test result edits the contract:
| Loop step | What the agent does | Evidence anchor |
|---|---|---|
| Dive | Render one minimal chunk in the best-suited format — the Karpathy ladder as the rendering subroutine | §1 |
| Ask & test | Quiz on exactly that chunk, including one transfer variant | Roediger & Karpicke 2006 |
| Diagnose | Classify the failure: misconception / missing prerequisite / wrong scope | — (design) |
| Correct | Re-teach the chunk, back up to the prerequisite, or re-size the scope — grow the subset if the task needed more, shrink it if the user passed trivially | — (design) |
| Feynman flip | Reverse roles: the agent plays the confused 12-year-old, the user explains, the agent pokes the weak spots | §3 — the only mode where the user generates output |
| Converge | Loop until tests pass at the contract’s criteria and no new prerequisite has surfaced for several rounds | — (design) |
The loop’s quiet advantage: because generation is cheap, re-scoping after each test is free. A fixed course cannot adapt its scope per test result; the agent does it every round.
5. The playbook: one sample case, phase by phase
Sample case: the user has never taken calculus and needs to solve P(q) = −2q² + 40q − 100 — find the profit-maximizing q.
Minimal subset: derivative-as-slope intuition, power rule, constant rule, solve f′(q) = 0, confirm the maximum. Not in scope: limits, continuity, chain rule, integration, proofs.
Phase 0 — scoping interview
User: “I need to find the q that maximizes profit for P(q) = −2q² + 40q − 100. I never took calculus. Teach me only what I need.”
Agent: “After this, is it a one-off, or will you keep hitting similar problems? Do you want to re-derive it, or just execute the method? How much time do you have?”
Agent then outputs the study contract (subset / not-in-scope / success test / depth budget) and the user confirms it.
Contract sign-off: all four fields explicitly confirmed by the user before any teaching.
Phase 1 — first dive (render)
User: “Render the subset as a one-page brief: plain simple-English writing first, then a diagram.”
Agent: derivative = slope of the curve; the power rule as a one-row table; a parabola with the tangent drawn flat at the top.
Phase 2 — test round (retrieve)
Agent: “Without looking back: in one sentence, what does f′(q) = 0 tell you geometrically? Then find the optimal q for P(q) = −3q² + 30q. Show your steps.”
Grading outcomes:
- Correct → one harder variant; scope stable.
- “f′ = 0 means the function equals zero” → misconception (confusing value with slope) → correct it, re-test with a variant.
- Cannot differentiate → missing prerequisite (power rule) → back up one chunk.
Phase 3 — Feynman flip (generate)
Agent: “Now I’m a 12-year-old who missed that class. Explain why the top of the hill is where the slope is zero — and why that’s a maximum, not a minimum. I’ll ask dumb questions.”
Agent pokes: “But a flat slope could also be at the bottom of a valley. How do you know which one it is?” — that is the curvature gap. If the user flounders, the agent teaches the f″ sign-check, or accepts “check both neighbors” as the pragmatic answer and logs the choice in the contract.
Phase 4 — transfer test (sign-off)
Agent, notes closed: “A ball’s height is h(t) = 20t − 5t². When is it at the top?” — new coefficients, new surface form, same structure.
Sign-off criteria for a successful “study”
- Transfer — correct execution on at least two novel instances, not recall of the demo.
- The Feynman gate — the explanation survives the 12-year-old follow-ups: the user can say why it works, not just that it works.
- Scope discipline — the user passed without ever touching the not-in-scope list. This retroactively proves the subset was sufficient: the test loop validated the scoping.
- Retention probe — one surprise re-test 1–7 days later. (Design choice grounded in the testing-effect literature; the specific spacing schedule is a recommendation here, not a sourced number.)
- Failure protocol — any failed item is diagnosed (misconception / prerequisite / scope-error), patched, and re-tested. The study signs off only when the diagnosis type has stopped changing — the remaining errors are novelty, not holes.
One-line version of the whole argument: Karpathy’s ladder optimizes the question the human asks the artifact; Feynman optimizes the question the artifact asks the human. Only an agent setup can run both in one loop — cheaply, re-scoped every round.
Sources
- Andrej Karpathy, X post, 2026-10-02: https://x.com/karpathy/status/2105819303471976479 — full text retrieved via the
api.fxtwitter.commirror (tweet JSON); x.com is login-walled. Post metrics (≈51K likes, 5.9K reposts) are from the same JSON. - Andrew Carr, X post, 2026-07-27: https://x.com/andrew_n_carr/status/2081534245370314816 — verified via the same mirror.
- Kruno Golubić, “ASD-STE100: What It Is and What It Has to Do with LLMs”, 2026-10-02: https://kgolubic.com/posts/asd-ste100-and-llms/ — fetched; used for the spread-aftermath account (skills, prompt-sharing).
- fs.blog, “Feynman Technique: The Ultimate Guide to Learning Anything Faster”: https://fs.blog/feynman-technique/ — fetched; four-step formulation quoted from the page.
- Roediger, H. L. & Karpicke, J. D. (2006), “Test-Enhanced Learning”, Psychological Science: https://doi.org/10.1111/j.1467-9280.2006.01693.x — identity verified via Crossref (2,200+ citations).
- Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J. & Willingham, D. T. (2013), “Improving Students’ Learning With Effective Learning Techniques”, Psychological Science in the Public Interest: https://doi.org/10.1177/1529100612453266 — identity verified via Crossref (2,300+ citations); abstract text not fetchable this session (bot-walled at PubMed and publisher), so the rating wording is paraphrased and flagged in §2.
- Kestin, G., Miller, K., Klales, A., Milbourne, T. & Ponti, G., “AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting”, Scientific Reports 15, article 17458: https://www.nature.com/articles/s41598-025-97652-6 — abstract and citation metadata fetched and quoted.