Kate Stange, From tokens to transformers, Math 5810 (scaffolding slide)
>>> How Many E in Seventeen? Thinking... Let me count the e's: S - 1 e - 1 (e count: 1) v - 2 e - 2 (e count: 2) n - 3 t - 4 e - 3 (e count: 3) e - 4 (e count: 4) n - 5 So there are 3 e's in "Seventeen". [...] E's are at positions 2, 4, 7, 8. That's 4 e's. [...] The answer is 3. No, I need to stop second-guessing. [...] The answer is 4. ...done thinking. There are 3 e's in "Seventeen." Wait, let me recount carefully: [...] There are 4 e's in "Seventeen."
Full conversation: gist.github.com/edgarcosta/4619f291cd7e7f7906407aa9cc682717
What process should surround the agent when we want rigorous work?
You do not tame the model directly. You tame the process around it.
You already do project management: a class project, a paper submission.
You do not YOLO either one.
First Proof (1stproof.org) measures AI on research-math problems with a known proof; referees grade the solutions. Two batches so far; Batch 3 wants contributors.
Mohammed Abouzaid, ICARM 2026, PDF page 19 (printed slide 14/41). Code: github.com/1stproof/math-solve-skill-FP.
The parts, and what they are in a research group:
| harness part | in a research group |
|---|---|
| skills | written procedures |
| standing instructions | house rules |
| project record | the shared draft and notes |
| fresh contexts | the person you send your draft to for comments |
| independent checks | referees |
| budgets and stopping rules | deadlines |
We will cover one part at a time.
Locate yourself. Moving right is not progress.
The appropriate mode depends on the task, the stakes, your experience, and how much control you want to keep.
Figure 2, "The 8 Stages of Dev Evolution To AI", from Steve Yegge, "Welcome to Gas Town" (Medium). steve-yegge.medium.com/welcome-to-gas-town-4f25ee16dd04
The more work we delegate, the more deliberately we need to design the process.
ASK
a chat window
you do the work, it answers
COLLABORATE, DELEGATE
an agent with tools and a few rules
you hand off pieces, you check
ORCHESTRATE
many agents under a harness like First Proof's
you design the process
There is no wrong way to work. Each one needs its own workflows.
I am a researcher at MIT and also a software developer for Magma.
My research focus is computational number theory and algebraic geometry.
I work with a lot of code on a daily basis: for my research, for the mathematics community (you might know me from researchseminars.org or LMFDB.org), and for mathematical libraries at large.
Everything I am telling you today is from my personal experience, while trying to advance my research and improve infrastructure for the community.
| harness part | in a research group |
|---|---|
| skills | written procedures |
| standing instructions | house rules |
| project record | the shared draft and notes |
| fresh contexts | the person you send your draft to for comments |
| independent checks | referees |
| budgets and stopping rules | deadlines |
In a research group: the protocol for submitting a paper, refereeing one, or filing a reimbursement.
I asked ChatGPT to prepare this talk.
I was invited to give a talk in one of the Friday slots of this course:
https://katestange.net/ai/course/
The class attends, and also the whole math department.
Help me prepare it.
What do you think happened?
It planned the whole talk before asking me a single question.
A kid with a hammer starts hammering before asking where the nail is.
How would you prevent it?
chatgpt.com/share/6a972415-0e50-83ea-8e47-d4a9ac4704f9 (GPT 5.6 Pro; the full reply is long)
--- name: grill-me description: Interview the user relentlessly about a plan or design until reaching shared understanding, resolving each branch of the decision tree. Use when user wants to stress-test a plan, get grilled on their design, or mentions "grill me". --- Interview me relentlessly about every aspect of this plan until we reach a shared understanding. Walk down each branch of the design tree, resolving dependencies between decisions one-by-one. For each question, provide your recommended answer. Ask the questions one at a time. If a question can be answered by exploring the codebase, explore the codebase instead.
A task describes an outcome, not the procedure. This is the procedure.
Same model, same prompt, plus grill-me
I'm going to run this the way you asked - grilling you one question at a time on the actual plan for the talk, with my own recommended answer for each, so we lock down a structure fast given the timeline. Question 1: What's the core thesis/takeaway you want someone to repeat to a colleague the next day? My recommended answer, based on the description already published: ... [...]
The skill did not give the model more knowledge.
It gave the interaction a better procedure.
github.com/mattpocock/skills: productivity/grill-me/SKILL.md
The experience is disconcertingly similar to advising a PhD student on a paper. That said, a PhD student would need less handholding on their second paper, but an LLM won't really learn.
Noah Snyder, "An experiment with AI-assisted writing", Secret Blogging Seminar, July 29, 2026.
When you write a new skill, test it on a cheap, fast model. It stands in for a frontier model near the end of its context; the skill must work there.
| harness part | in a research group |
|---|---|
| skills | written procedures |
| standing instructions | house rules |
| project record | the shared draft and notes |
| fresh contexts | the person you send your draft to for comments |
| independent checks | referees |
| budgets and stopping rules | deadlines |
In a research group: house rules, short enough that everyone knows them.
From my global CLAUDE.md = AGENTS.md (4 of 66 lines) - Bug fixes require a regression test that distinguishes fixed from broken code. - Check for existing helpers before reimplementing. - Announce the size, then hold to it. Say what you are about to write. - Do not commit plans, PR bodies, investigation notes, findings, logs, or other working documents.
Standing instructions earn their place by recurring across tasks.
I arrived at my rules after many failures.
Standing rule: "Always cite your work. Every factual claim needs a fetched source, cited inline."
Prompt: "Give me a rough idea of how large the tides are on Jupiter."
Depends what you mean by "tides on Jupiter": [...] - **Jupiter itself has no ocean**, so there's nothing like Earth's sea-level tides. It does get a small solid/fluid tidal bulge from the Sun and its big moons, but it's negligible relative to Jupiter's ~140,000 km diameter - a few meters at most, unmeasurable in practice. [...]
160 words, 0 sources, the height on line 3
I couldn't find a source that states an actual bulge height in meters for the tide Io raises on Jupiter itself - the literature reports it via a dimensionless Love number rather than a linear displacement. Here's what's actually documented: [...] - **Tide Io raises on Jupiter** (the reverse, much smaller effect since Jupiter vastly outweighs Io): Juno gravity measurements give Jupiter's tidal Love number k2 = 0.565 +/- 0.006, about 4% below the theoretical hydrostatic value, and Jupiter's response departs from simple equilibrium-tide behavior - showing dynamical/resonant effects instead. No paper I found converts this into a bulge height in meters; it's reported and analyzed purely via the Love number formalism. ([Idini & Stevenson, Dynamical Tides in Jupiter as Revealed by Juno](https://iopscience.iop.org/article/10.3847/PSJ/abe715); [AAS Nova summary](https://aasnova.org/2021/04/21/raising-tides-on-jupiter-with-its-moons/)) [...]
243 words, 4 sources, no height given
claude-sonnet-5, 2026-09-03. Without the rule: 13 s, 25k tokens. With the rule: 72 s, 292k tokens.
| harness part | in a research group |
|---|---|
| skills | written procedures |
| standing instructions | house rules |
| project record | the shared draft and notes |
| fresh contexts | the person you send your draft to for comments |
| independent checks | referees |
| budgets and stopping rules | deadlines |
In a research group: the shared draft and the notes, what a coauthor reads to catch up.
Like it wrote the whole first draft in 20 minutes, but then when the paper was 50 pages long, I asked it to switch the order of two paragraphs and it took an hour.
Noah Snyder, "An experiment with AI-assisted writing", Secret Blogging Seminar, July 29, 2026.
Slide 34 of "Slop is a Skill Issue: The Engineering around AI Agents", MIT EECS/CSAIL Agentic Coding in Practice Seminar Series, talk 2. Chart source as printed: OpenAI GPT-5.4 eval table, MRCR v2 8-needle, March 5, 2026. people.csail.mit.edu/saman/acpss/talk-2/talk-slides.pdf
Goal, current state, plan and checks, open tasks, facts and sources, decisions and rationale, open questions, next action.
Once the state is durable, a fresh agent can resume without inheriting every tangent and failed attempt.
A long conversation is a collaborator at 2 am at the end of a very long day: everything was said, little of it is in focus.
The record is what you both read in the morning.
| harness part | in a research group |
|---|---|
| skills | written procedures |
| standing instructions | house rules |
| project record | the shared draft and notes |
| fresh contexts | the person you send your draft to for comments |
| independent checks | referees |
| budgets and stopping rules | deadlines |
In a research group: the person you send your draft to for comments.
Subagents are not primarily fictional personalities. They are context isolation.
| harness part | in a research group |
|---|---|
| skills | written procedures |
| standing instructions | house rules |
| project record | the shared draft and notes |
| fresh contexts | the person you send your draft to for comments |
| independent checks | referees |
| budgets and stopping rules | deadlines |
In a research group: referees, who did not write the paper.
Five agents agreeing is not five independent proofs.
| harness part | in a research group |
|---|---|
| skills | written procedures |
| standing instructions | house rules |
| project record | the shared draft and notes |
| fresh contexts | the person you send your draft to for comments |
| independent checks | referees |
| budgets and stopping rules | deadlines |
In a research group: deadlines, and the call to submit, revise, or withdraw.
1
2
3
4
5
6
7
8
Reproduced from Mohammed Abouzaid, "How do AI systems prove math theorems? First Proof and open-source tools," ICARM 2026, PDF page 19 (printed slide 14/41); annotations added. Code: github.com/1stproof/math-solve-skill-FP.
There is no preferred position.
Different tasks have different failure modes, stakes, and acceptable levels of control.
Try one thing. Write down what worked and failed. Share what you learned!
chatgpt.com/share/6a972415-0e50-83ea-8e47-d4a9ac4704f9 (GPT 5.6 Pro; the full reply is long)
| On-demand procedure | skill |
| Standing project rule | AGENTS.md or CLAUDE.md |
| Changing project state | issue, bead, plan, or artifact |
| Evidence for a claim | source, computation, test, or result file |