Coaching the adult who teaches, rather than tutoring the child
Every AI tool for home education I could find either plans the week ahead or teaches the child directly, and both leave the adult doing the teaching to work it out alone. Paideia is the record of what actually happened in a lesson, turned into coaching for the person who taught it, designed and built solo and running with one pilot household.
- The turn. I built a way to read her handwritten planner, and it worked, but it read the wrong thing. A pilot “no” on 28 July moved the build from approving lessons to being asked about them. On 2 September she said “this is the whole thing.”
- The state. Solo, one household, four moderated sessions, still running. The most recent evidence cuts both ways, and the page says how.
Short on time? The timeline below and “What this shows about how I work” near the end are the quick version.
- Role
- Solo, covering research, design and build
- Built with
- Expo, React Native, Supabase, Claude API
- When
- 2026, from a proof of concept in February to one pilot household and four moderated sessions across ten weeks
- Results
-
- 34 handwritten planner pages captured in 11 minutes, with no instruction
- 4 errors in 51 cells checked against a hand-built reference set, and no corrections across 425 cells she reviewed on 24 of 28 pages
One household, my own. Each moderated session’s accept criteria were written down beforehand.
- Key decision
- After the pilot said no, I replaced the review queue with a conversation. Her own words can’t be misread the way handwriting can, so nothing she tells it waits for approval.
Context
My wife home-educates our children, and for two years she had been logging each day's lessons in a paper planner. It held almost everything she knew about how the teaching was going, and nothing ever read it back to her.
Paideia started in February as a Claude project. I added transcripts of recorded lessons and asked whether any of it would help the person teaching. It did, which turned the question into how to make that a product, and that household became its only pilot.
The problem
Home educating is not a smaller version of school. There is no department to compare notes with, no scheme of work to measure against, and no external mark of whether a year went well. The tools that exist mostly assume otherwise.
- Working alone. No colleague to think out loud with, and in a home-educating context admitting that a lesson went badly carries a weight it would not carry in a staffroom.
- A record that only records. The planner captured each day faithfully and gave nothing back.
- No view of the arc. Nothing showed how a child had moved across a term or a year, which matters more now that the Children's Wellbeing and Schools Act requires every local authority to keep a register of children not in school.
How the product changed its mind
Five moments from the pilot, in order.
-
Feb
Recorded lessons, read in a Claude project
Recorded lessons, five taught by me and three by her, read in a Claude project to see whether any of it would help the person teaching. The question became how to make it a product.
-
Jun
Planner photo to structured lessons
A photograph of her handwritten planner becomes structured lessons. The extraction worked, and turned out to be reading the wrong thing.
-
Jul
She said no
Asked whether photographing the planner was worth it, she said no. The cost was checking every lesson. Accept criteria were written down beforehand.
-
Aug–Sep
Capture became a conversation
What she tells it is never queued for approval, and a conversation of length one is a complete record.
-
Sep
Themes, then ideas for when she is stuck
On 2 September she called a month-level view “the whole thing”. Ideas for when she is stuck, each tied to where it came from, are the build that followed.
Research and insight
The tools split cleanly into two bands. Planners point forward at the week you intend to teach, and direct-to-child tutors replace the adult altogether. Neither does anything for the person doing the teaching.
-
Planners
Point forward
At the week you intend to teach.
-
Direct-to-child tutors
Replace the adult
The child is taught directly.
-
Paideia
Coaches the adult
Reads what actually happened in a lesson, and coaches the person who taught it.
Why I thought the gap was real
What convinced me the space between them was real rather than merely empty was Erik Hoel's essay on why we stopped making Einsteins, which argues that one-to-one tutoring is the only method that has reliably produced them, and that it vanished because it cannot be standardised. Einstein had Max Talmud. The practice is old and reasonably well evidenced. It had simply never been productised.
The February recordings added the finding the rest of the project kept returning to. Whether a child got an answer right says little on its own, because the useful part is the route there: how much prompting it took, which object on the table helped, and whether it held once the support was taken away.
Strategic reframe
She was already logging retrospectively, and using that logging as her reflection. Designing forward would have meant competing with planners and asserting what a week ought to contain, which in home education is the wrong posture, because the approach is tailored to the child by definition. So the tool reads what actually happened and coaches from that, which makes it empirical rather than prescriptive.
The answer to whether the grandmaster or the computer is stronger is the grandmaster holding the computer, and Paideia makes the same bet about the person teaching.
Extraction before screens
The whole product rested on one question, whether a photograph of a handwritten planner could become structured lessons, so I answered it before designing anything. In May the extraction read a real planner page with more than 90% accuracy, checked by hand, and the screens followed in June.
How I measured the error rate
Measuring it properly took longer. A check that runs the model several times only measures agreement, and a confident consensus on a misread digit is the most dangerous thing this system can produce, so I built a reference set by hand. Across 51 cells it found 4 errors, or 7.8%.
Her own corrections gave the wider picture. 24 of 28 pages needed none across 425 cells, and the 4 that failed, at 27 to 43%, were all the right-hand page of a spread. I had put it down to volume, and it was a missing header: those pages carry no child names.
The extraction worked, and it turned out to be reading the wrong thing. A planner records what was taught and cannot hold how the child did. As she put it in the second session, a line saying a child did three pages of work cannot tell the system whether they were three essays or three paragraphs.
The no, and what earned the yes
Each moderated session had accept criteria written down beforehand, so a disappointing result would be a finding rather than a matter of interpretation. The third, on 28 July, asked whether photographing her planner had been worth doing, and the answer was no.
“Because it’s two children and it’s every day, it’s quite a lot of information to process… I have to cross-reference with my planner, which feels quite tedious.”
Pilot tutor, 28 July 2026It was not a rejection of the idea. In the same session she described what she would value, a picture of the weeks rather than a card after every lesson and quick ideas for when she was stuck, and she named what cost her most, which was approving every lesson one at a time.
So that is what I built next. On 2 September she was shown themes and patterns across a month for each child and subject, with a way to consolidate or expand any topic in it, and the answer arrived before the question was asked.
“This is what I feel like the app should be. Rather than just record keeping, it’s like this is the whole thing.”
Pilot tutor, 2 September 2026The same session said what was still wrong. The part she valued sat five taps from the home screen, and the synthesis spent its words handing her own facts back, which left her fact-checking it before she could trust the next line. Her fix became a rule: facts she gave come back as short lines with their source, and prose is spent only on what to try next.
Looking for the right input
Better synthesis could not fix the deeper problem, which she named in August when a summary felt like it was handing back what she already knew. Every source in the record was her own account, and a system fed only one person's account can enrich it but never surprise it, so the search turned to evidence she had not written.
What else I tried and why it failed
Each candidate failed on something different. Photos of worksheets were the hardest to read: in the first trial four of six batches were lost to image rotation before extraction was reached, and all three submissions were filed against the wrong child. Recordings of whole lessons were the richest source and the noisiest, because a kitchen table is not a classroom, and one September recording held two lessons on a single microphone with a three-year-old in the room. Recorded voice notes would have meant storing a family's audio, with the retention and consent questions that brings, which was more than a pilot of one could justify.
What worked better was asking. At the end of August capture became a conversation rather than a form, because a form receives what it is given and a conversation can notice what is missing. Her first real session with it showed that being asked read as interest rather than appraisal, so the next build has it ask, after each lesson, where things stalled and what she had to re-explain, and offer coaching at the end rather than volunteer it.
That does not resolve the tension, it gives it somewhere to sit. What she tells it is reliable about her read of a lesson and can only hold what she noticed, while a recording or a worksheet can show her something new but first needs her to say who was in the room. The conversation is where the two meet, asking only the one or two cheap questions that decide what the other evidence is allowed to say.
The conversation also changed where the value sat. The synthesis she had called the whole thing lived in a different part of the app from the place she spent her time, and reached her only after a wait and a trip between the two. A chat that answered in seconds, on the same screen, collapsed that, and left the synthesis as somewhere to look back rather than the place the value lived.
The same format changed what there was to act on. A long record of logged lessons had often yielded nothing to suggest, because each described what happened rather than how it went, while one conversation about one lesson could now yield several things to try.
Design decisions
Most of the decisions that mattered were about restraint rather than capability. The model can do considerably more than the product lets it do, and almost every rule below exists to stop it doing something it would otherwise do fluently and wrongly.
A teaching assistant, not a teacher
A tool that tells a home educator what she ought to be teaching is a red rag, because not being told what to do is often part of why a family is home educating in the first place. So Paideia sits as a teaching assistant rather than a teacher, a coach or a mentor. It can notice things and ask about them, and it can suggest a way to consolidate something she has taught, but it never decides what comes next.
The cost was the checking
The pilot showed the real cost was checking, because she had to approve every lesson against her planner, and she asked to see only the ones the system was unsure of. So lessons read with high confidence could be accepted in one opt-in batch, and the rest queued for her with the reason attached.
It helped less than it should have. On 3 August she submitted 34 pages in 11 minutes with no instruction and still went through them card by card. The conversation removed the queue instead, because her own words cannot be misread the way handwriting can, so nothing she tells it waits for approval.
A conversation of length one is complete
A surface which can ask can drift into one that expects an answer, so the floor is explicit: she can open it, drop in a photograph, say nothing and leave, and that is a complete record. Questions are offered and never gating. I track the proportion of one-line conversations, because if it ever reaches zero the floor has stopped being real.
A bare loading spinner is a bug
Three eight-second waits in one sitting are not remembered as twenty-four seconds, they are remembered as a tool that always makes you wait. So every wait is either removed from her path, hidden behind something she was going to do anyway, or filled with an honest account of the work actually happening. That last option is the most tempting and the most dangerous, because the first time an animation is caught overstating what it is doing, every honest one after it reads as theatre.
The record has no prose field
An early version praised a child's progress on a fortnight of thin data, and it read as flattery, which cost more than saying nothing would have. The fix was structural rather than a change of tone. The generated record has no free-text field at all: it is assembled from coverage, typed observations and quotes copied out by code, and every entry carries a reference back to the thing it came from. Where her account and the worksheet disagree, both go in, unresolved. The coaching is written as prose, because it is a suggestion she can take or leave rather than a claim about her child.
Building it
I built it solo. Expo and React Native on the front, Supabase for data and auth, the Claude API for extraction and synthesis, deployed as a progressive web app so it installs on a phone without an app store. The hard part was never the interface, it was the schema. A worksheet page is not one piece of evidence, it is twelve, and any claim the tool makes months later has to be able to point back at the right one.
The habit that paid for itself was checking what the system actually wrote rather than what it reported. The day she submitted 34 pages, 10 of the submissions reported success and wrote nothing, from three separate causes, and fixing them recovered 230 lessons. A day later a refactor broke the live path while every test passed, because every test ran without writing, so now any change that writes to her record is run once for real before I call it done.
Constraints and trade-offs
Two things I decided not to build. The first is curriculum advice. Telling a tutor what a child ought to learn next makes the tool the authority, and the premise of the whole thing is that the tutor is.
The second is anything touching additional needs. I talked it through with a child educational psychologist I know and concluded that the distance between noticing a pattern and implying a diagnosis is shorter than it looks, and not a line for a tool at this stage to walk up to.
The pilot is one household, so the findings are deep rather than broad, and the household is mine, so her feedback came to me as her husband. Writing the accept criteria down before each session, and keeping the no when it came, is how I tried to stop that softening the evidence.
Where it stands
There are no commercial outcomes here. There is one household, four moderated sessions across ten weeks, and a pilot that is still running.
That first real conversation, on 21 September, is the most recent evidence and it cuts both ways. The coaching was grounded in her own records and she took most of it, but she also said she had no real interest in it creating a record, and that she would not use it again if she had downloaded it, because it had told her three times that it had saved lessons it never saved. Fixing that comes before asking her for anything more.
Before the current pilot window opened she recorded what she wants the app to produce. The window closes in mid-October, and I will judge it against that list rather than against my own sense of how it went.
Managing designers, I could always describe a constraint and hand it over. Building this alone meant every decision I made about posture had to survive contact with a schema, and a few of them did not.
What this shows about how I work
- I listen for where the value lands. I set out to build a long view of a child; she felt the value in everyday use. In July she wouldn’t pay for a tool like this; by September it was “the whole thing”. Felt value sits upstream of revenue, so I start there.
- Evidence over opinion. I wrote the accept criteria down before each session, so a disappointing result would be a finding rather than a matter of interpretation, and I kept the “no” when it came.
- Restraint as a design decision. The model can do more than the product lets it. I decided where it stops: it assists rather than teaches, it never goes near a diagnosis, and the record it produces has no free-text field to flatter with.
- I check what the system wrote, not what it reported. The day she submitted 34 pages, ten submissions reported success and wrote nothing. Fixing that recovered 230 lessons, and any change that writes to her record is now run once for real before I call it done.
Reflections
The rule that the record can only say what the evidence supports sounds like a values statement right up until you are the one writing the table that enforces it, at which point it becomes a column and a migration.
The other thing I had not expected to spend so much time on was the law. Holding information about a named child is legally fraught, and working through it turned what had been a design preference into the thing that makes the product defensible: Paideia does not assess the child, it supports the adult teaching them. It records what happened in a lesson, but every output is aimed at the tutor's practice rather than at a verdict about her child.
What is unresolved is whether any of it generalises. One household produced findings I trust and a sample I do not, and the only way to know which parts are about this family and which are about home education is to put it in front of households I am not part of.