Kate Stange, From tokens to transformers, Math 5810 (scaffolding slide)
>>> How Many E in Seventeen? Thinking... Let me spell it out: S-e-v-e-n-t-e-e-n Let me count the e's: S - 1 e - 1 (e count: 1) v - 2 e - 2 (e count: 2) n - 3 t - 4 e - 3 (e count: 3) e - 4 (e count: 4) n - 5 So there are 3 e's in "Seventeen".
Wait, let me recheck: S-e-v-e-n-t-e-e-n [...] E's are at positions 2, 4, 7, 8. That's 4 e's. [...] Hmm, but wait. Let me be very careful. [...] The answer is 3. No, I need to stop second-guessing. [...] The answer is 4. ...done thinking. There are 3 e's in "Seventeen." Wait, let me recount carefully: [...] There are 4 e's in "Seventeen."
Full conversation: gist.github.com/edgarcosta/faddc970054c8ba53afbbcf25fc916e8
What process should surround the agent when we want rigorous work?
You do not tame the model directly. You tame the process around it.
You already do project management: collaborative work, writing a paper.
You do not YOLO either one.
First Proof (1stproof.org) measures AI on research-math problems with a known proof; referees grade the solutions. Two batches so far; Batch 3 wants contributors.
Mohammed Abouzaid, ICARM 2026, PDF page 19 (printed slide 14/41). Code: github.com/1stproof/math-solve-skill-FP.
The parts, and what they are in a research group:
| harness part | in a research group |
|---|---|
| skills | written procedures |
| standing instructions | house rules |
| project record | the shared draft and notes |
| fresh contexts | the person you send your draft to for comments |
| independent checks | referees |
| budgets and stopping rules | deadlines |
We will cover one part at a time, using analogies.
How do you interact with models? Moving right is not necessarily progress.
Figure 2, "The 8 Stages of Dev Evolution To AI", from Steve Yegge, "Welcome to Gas Town" (Medium). steve-yegge.medium.com/welcome-to-gas-town-4f25ee16dd04
The more work we delegate, the more deliberately we need to design the process.
ASK
a chat window
you do the work, it answers
COLLABORATE, DELEGATE
an agent with tools and a few rules
you hand off pieces, you check
ORCHESTRATE
many agents under a harness
you design the process
There is no wrong way to work. Each one needs its own workflows.
I am a researcher at MIT and also a software developer for Magma.
My research focus is computational number theory and algebraic geometry.
I work with a lot of code on a daily basis, for:
Everything I am telling you today is from my personal experience, while trying to advance my research and improve mathematical infrastructure.
| harness part | in a research group |
|---|---|
| skills | written procedures |
| standing instructions | house rules |
| project record | the shared draft and notes |
| fresh contexts | the person you send your draft to for comments |
| independent checks | referees |
| budgets and stopping rules | deadlines |
In a research group: the protocol for writing a paper, refereeing one, or filing a reimbursement.
I asked GPT 5.6 Pro. It worked for twelve minutes.
I was invited to give a talk in one of the Friday slots of this course:
https://katestange.net/ai/course/
The class attends, and also the whole math department.
Help me prepare it.
What do you think happened?
It planned the whole talk before asking me a single question.
A kid with a hammer starts hammering before asking where the nail is.
How would you prevent it?
chatgpt.com/share/6a972415-0e50-83ea-8e47-d4a9ac4704f9 (GPT 5.6 Pro; the full reply is long)
--- name: grill-me description: Interview the user relentlessly about a plan or design until reaching shared understanding, resolving each branch of the decision tree. Use when user wants to stress-test a plan, get grilled on their design, or mentions "grill me". --- Interview me relentlessly about every aspect of this plan until we reach a shared understanding. Walk down each branch of the design tree, resolving dependencies between decisions one-by-one. For each question, provide your recommended answer. Ask the questions one at a time. If a question can be answered by exploring the codebase [the project files], explore the codebase instead.
Same prompt, plus grill-me (Claude Sonnet, fresh conversation)
I'm going to run this the way you asked - grilling you one question at a time on the actual plan for the talk, with my own recommended answer for each, so we lock down a structure fast given the timeline. Question 1: What's the core thesis/takeaway you want someone to repeat to a colleague the next day? My recommended answer, based on the description already published: ... [...]
The skill did not give the model more knowledge.
It gave the interaction a better procedure.
github.com/mattpocock/skills: productivity/grill-me/SKILL.md
Noah Snyder on "An experiment with AI-assisted writing", Secret Blogging Seminar, reported:
The experience is disconcertingly similar to advising a PhD student on a paper. That said, a PhD student would need less handholding on their second paper, but an LLM won't really learn.
A way to test how comprehensible a skill is, is to test it on cheap/fast models.
You are the one who learns; the skill carries the lesson into the next conversation.
| harness part | in a research group |
|---|---|
| skills | written procedures |
| standing instructions | house rules |
| project record | the shared draft and notes |
| fresh contexts | the person you send your draft to for comments |
| independent checks | referees |
| budgets and stopping rules | deadlines |
In a research group: house rules, short enough that everyone knows them.
From my standing-instructions file, CLAUDE.md = AGENTS.md (7 of 66 lines) - Bug fixes require a regression test [a test that fails on the old code] that distinguishes fixed from broken code. - Check for existing helpers [library functions] before reimplementing. - Do not comment idiomatic patterns [routine code]. - Be concise in PR bodies [change descriptions], code comments, and tracker comments. - Announce the size, then hold to it. Say what you are about to write. - Do not commit [save into the shared history] plans, PR bodies, investigation notes, findings, logs, or other working documents. - Use Conventional Commits [a fixed format for change summaries]: <type>(<scope>): <description>, scope optional.
I arrived at my rules the hard way.
Standing rule: "Always cite your work. Every factual claim needs a fetched source, cited inline."
Prompt: "Give me a rough idea of how large the tides are on Jupiter."
Depends what you mean by "tides on Jupiter": [...] - Jupiter itself has no ocean, so there's nothing like Earth's sea-level tides. It does get a small solid/fluid tidal bulge from the Sun and its big moons, but it's negligible relative to Jupiter's ~140,000 km diameter - a few meters at most, unmeasurable in practice. [...]
160 words, 0 sources, 25k tokens, the height on line 3
I couldn't find a source that states an actual bulge height in meters for the tide Io raises on Jupiter itself - the literature reports it via a dimensionless Love number rather than a linear displacement. Here's what's actually documented: [...] - Tide Io raises on Jupiter (the reverse, much smaller effect since Jupiter vastly outweighs Io): Juno gravity measurements give Jupiter's tidal Love number k2 = 0.565 +/- 0.006, about 4% below the theoretical hydrostatic value, and Jupiter's response departs from simple equilibrium-tide behavior - showing dynamical/resonant effects instead. No paper I found converts this into a bulge height in meters; it's reported and analyzed purely via the Love number formalism. ([Idini & Stevenson, Dynamical Tides in Jupiter as Revealed by Juno](https://iopscience.iop.org/article/10.3847/PSJ/abe715); [AAS Nova summary](https://aasnova.org/2021/04/21/raising-tides-on-jupiter-with-its-moons/)) [...]
243 words, 4 sources, 292k tokens, never gave the requested rough scale
Without the rule: 13 s. With the rule: 72 s, 12 times the cost.
| harness part | in a research group |
|---|---|
| skills | written procedures |
| standing instructions | house rules |
| project record | the shared draft and notes |
| fresh contexts | the person you send your draft to for comments |
| independent checks | referees |
| budgets and stopping rules | deadlines |
In a research group: the shared draft and the notes, what a coauthor reads to catch up.
Noah Snyder also observed:
Like it wrote the whole first draft in 20 minutes, but then when the paper was 50 pages long, I asked it to switch the order of two paragraphs and it took an hour.
One model's accuracy at recalling one fact from the conversation: 97% when short, 37% when very long.
A long conversation is a collaborator at 2 am at the end of a very long day: everything was said, little of it is in focus.
The record is what you both read in the morning.
Slides 34 and 37, "Slop is a Skill Issue" (MIT), people.csail.mit.edu/saman/acpss/talk-2/talk-slides.pdf
This allows one to use subagents that are ready to go.
Agents take as fact whatever they read, their own guesses included; label each claim by its support.
| harness part | in a research group |
|---|---|
| skills | written procedures |
| standing instructions | house rules |
| project record | the shared draft and notes |
| fresh contexts | the person you send your draft to for comments |
| independent checks | referees |
| budgets and stopping rules | deadlines |
In a research group: the person you send your draft to for comments.
| harness part | in a research group |
|---|---|
| skills | written procedures |
| standing instructions | house rules |
| project record | the shared draft and notes |
| fresh contexts | the person you send your draft to for comments |
| independent checks | referees |
| budgets and stopping rules | deadlines |
In a research group: referees, who did not write the paper.
Five agents agreeing is not five independent proofs.
| harness part | in a research group |
|---|---|
| skills | written procedures |
| standing instructions | house rules |
| project record | the shared draft and notes |
| fresh contexts | the person you send your draft to for comments |
| independent checks | referees |
| budgets and stopping rules | deadlines |
In a research group: deadlines, and the call to submit, revise, or withdraw.
Every box in First Proof's diagram is one of the six parts you just saw; yours can be twenty lines.
Different tasks carry different pitfalls and stakes, and deserve different oversight.
Try one thing. Write down what worked and what did not. Share what you learned!
1
2
3
4
5
6
7
8
Figure: M. Abouzaid, ICARM 2026, slide 14; code: github.com/1stproof/math-solve-skill-FP