We tried 31 of them on twelve real university lectures (MIT math and computer science, Yale philosophy and history) and checked every set of notes against answers we'd worked out by hand first. Nine were good enough to put on this page. Here's the one to pick, and what it costs.
Notes matter most when the material is brand new, and that's exactly when a wrong fact is hardest to catch. Nobody can spot an error in something they're learning for the first time; that's just what learning is. A confidently-wrong note gets studied and trusted, which makes it worse than a note that's thin. So our number-one rule, before anything else, is: never make something up.
That's why our testing isn't about which AI sounds smartest. It's about which one we can trust not to invent a number, misattribute a quote, or botch a calculation; then, among the trustworthy ones, which keeps the most useful detail.
All three produce the same kind of notes; they differ in setup, privacy, and ceiling.
The option LectureSync ships with. Runs on your Mac using Apple's built-in models. Zero setup.
Run a custom open model on your own Mac with Ollama, LM Studio, or oMLX, and point LectureSync at it.
Use a hosted model through a provider like OpenRouter. The highest quality ceiling, for cents a semester.
We used real lectures, not quiz questions. Twelve of them: MIT linear algebra, computer science and economics, plus Yale philosophy, biology and history. Hard science and heavy reading, because your timetable has both.
We marked the notes against our own answers. We sat through each lecture first and wrote down what should end up in the notes (the worked examples, the numbers, the names, the quotes), then checked what each model actually kept. We never asked another AI to grade the work.
Then the test that decides everything. One math lecture works out the inverse of a matrix, and the answer is a specific grid of numbers. If a model guesses and gets it wrong, you have no way of knowing: you're seeing the material for the first time, which is the whole reason you're reading the notes. So we asked every model that same question 300 times and checked all 300 answers by hand.
Six models got it wrong often enough that we dropped them, however good the rest of their notes were. Everything in the table below passed.
All nine passed the make-things-up test, so you can pick any of them and be safe. What's left to choose on is how much of the lecture they keep, how fast they are, and what they cost. The top row is the one we'd pick. Click any heading to re-sort.
| Our takeOur take | |||||
|---|---|---|---|---|---|
| Hunyuan 3 app pick | Tencent | Most detail of anything we tried, and it never slipped once in 300 tests | $0.32 | 1m 23s | |
| MiniMax M3app pick | MiniMax | Nearly as detailed as our pick, if you'd rather not use it | $0.48 | 1m 29s | |
| Ling 3.0 Flashapp pick | inclusionAI | The cheapest one we'd recommend, and it's quick | $0.07‡ | 9s | |
| DeepSeek V4 Flash 0731app pick | DeepSeek | The steadiest: it gave us the same quality note every time | $0.16 | 1m 34s | |
| GPT-5.6 Lunaapp pick | OpenAI | A good all-rounder: fast, detailed, a familiar name | $0.19 | 18s | |
| Qwen3.6 35B A3B | Alibaba | Fine, but it drops more of the specifics | $0.59 | 38s | |
| DeepSeek V4 Pro | DeepSeek | Costs the most on this page and keeps less | $0.72 | 1m 17s | |
| Gemini 3.5 Flash Lite | By far the fastest: notes in five seconds | $0.43 | 5s | ||
| Mistral Small 2603 | Mistral | Cheap and quick, but the notes come out thin | $0.15 | 8s |
The five shaded rows are built into LectureSync: you pick them from a menu, no typing.
Detail it keeps: how much of the real lecture survives into your notes: the worked examples, the numbers, the names.
Time per lecture: the writing happens in the background while you get on with your day, so the five-second one and the minute-and-a-half one feel the same.
Cost for a semester: 75 lectures, five a week for fifteen weeks. A five-dollar top-up covers years.
‡ Ling 3.0 Flash is free right now. We've listed what its maker charges everyone else, so the price on this page doesn't become a lie when the free run ends.
Three models we used to recommend aren't here any more. When we started asking 300 times instead of three, all three turned out to invent answers, not often, but often enough. A note with a confidently wrong number in it is the one thing we won't hand a student, so they're out.
| Model | How often it made something up | Verdict |
|---|---|---|
| Nemotron 3 Super NVIDIA | About 1 note in every 60 | Out |
| MiMo v2.5 Pro Xiaomi | About 1 note in every 150 | Out |
| MiMo v2.5 Xiaomi | About 1 note in every 150 | Out |
Our cut-off is one note in 200. Anything worse and we won't put it in front of you, however good the writing is. All three of these looked perfectly clean when we only asked three times, which is exactly why we stopped doing that.
We used to ask each model once and see how it did. That tells you how well something writes. It tells you nothing about the mistake it makes one time in a hundred, and one time in a hundred is once a term for you. So this round we asked 300 times each. Six models failed; three of them had been on this page until today.
One of those six wrote the best notes we saw all summer: more of the examples and numbers kept than anything else we tried. It also copies down one of the lecturer's numbers wrong about once every 150 notes, and then works from the wrong one. You would never catch it, because you're reading those notes to learn the material in the first place. Great writing doesn't get a model onto this page. Getting it right does.
Every solid dot is a model that passed. Higher is better, left is cheaper, so the best place to be is the top-left corner, and that's where the ones we recommend are. The two famous names on the right cost thirty times more.
Across: what a semester costs (the scale is squashed, or everything cheap would pile up on the left edge). Up: how much of the lecture survives. ⭐ = our pick. The two faded dots were only ever tested three times, not 300, which is why they're not on the list.
Three models on our old list looked perfect when we asked three times. Asked 300 times, all three made things up. Once in 60 notes for the worst one, which sounds tiny until it's the note you revise from, on a subject you're meeting for the first time. Nobody spots a mistake in something they're still learning.
Our pick costs 32¢ a semester and kept more of the lecture than anything else. On the table above, the cheapest model kept 89% of the details and the dearest kept 80%. You are not getting better notes by paying more here.
Once they're all honest, the question is what actually makes it into the notes: the worked example, the exact figure, the name of the theorist you'll be asked about. The gap is real: the best keeps 92% of that, the thinnest 60%. The week before an exam, that difference is the whole thing.
If you just want to be told, it's the first one. If you want to know why we landed there, here it is.
It kept more of the lecture than anything else we tried, and it's the only model in the whole test that came back perfect: we asked it the hardest question 300 times and it got it right 300 times. It takes about a minute and a half per lecture, which you'll never notice, because it writes while you're doing something else.
Seven cents for the whole term, notes in nine seconds, and it still keeps 89% of the detail. It's free at the moment; we've quoted the normal price so you're not surprised later. If the pennies matter more to you than the last few percent, take this one.
GPT-5.5 and Claude Opus 4.8 are excellent, and if you already pay for one you'll get lovely notes from it. But they cost $11.49 and $10.03 a semester next to 32¢ for our pick, and four cheaper models keep just as much of the lecture. That's not a trade we'd ask a student to make, so we stopped re-testing them.
The very fast, very cheap ones (Gemini 3.5 Flash Lite, Mistral Small 2603) finish a lecture in five and eight seconds. They also throw a lot away: 72% and 60% of the detail kept, against 92% at the top. They're on the table so you can see the trade, not because we'd choose them.
Very. We ran all of this on August 2, 2026. Five of the nine models on this list were less than two months old when we tested them, and one was two days old. New models turn up constantly, so we do the whole thing again when they do, including re-testing the ones that already passed, because that's where the nasty surprises come from.
Each dot is the day a model arrived on OpenRouter. All nine landed within five months of our August 2 test. The starred dot is Hunyuan 3, our pick.
Before the cloud study, we ran the same kind of bake-off on 13 local models: 1,755 scored generations on real MIT lectures, running entirely on-device. When Google shipped QAT checkpoints of the Gemma-4 family, we re-ran the accuracy tests with the same method; the picks below come from that re-test. Want every number? See the full local leaderboard: all 35 models we've tested, every score, and the rejects.
Q4_K_XL · 2.6 GBQ4_K_XL · 2.6 GBQ4_K_XL · MoE · 14.3 GBQ4_K_XL · MoE · 14.3 GBThe surprise of the original study held up in the re-test: the smallest model is the most trustworthy. E2B's QAT checkpoint invented a false statement just 0.3% of the time at half its old size, and the 26B mixture-of-experts came back with zero. That's why E2B is the pick on every Mac under 24 GB, not the budget fallback.
OpenRouter is one account that gives you access to almost every AI model out there. Instead of signing up with five different companies, you sign up once and pick models like items on a menu. That's why it's the easiest on-ramp (and what we used for this entire test).
Sign up like any website. Add five dollars of credit; at 32¢ a semester for our pick, that outlasts your degree.
In OpenRouter's settings, create a key. It's a long string starting with sk-or-…. Treat it like a password and don't share it.
Settings → Connections → add OpenRouter and paste your key. It's stored in your Mac's Keychain; LectureSync never sends it anywhere except OpenRouter itself.
Settings → Notes: for notes, our recommendation is tencent/hy3. For transcription, openai/whisper-large-v3-turbo is the standard pick (this bake-off covered the notes step; transcription wasn't part of it).
Step 3: your key, saved in Settings → Connections.
What "Measured" means here: every pick on this page comes from our own testing, marked against answer keys we wrote by hand. We don't ask another AI what it thinks of the notes.
Where your lecture goes: with the built-in or local options, your audio and notes never leave your Mac. Pick one of the models on this page and the text of your lecture goes to that company, under their terms. Details in our privacy policy.
Nobody pays to be on this page. New models land every month and we re-test. If a pick here changes, it's because something measured better, never because someone bought the spot.
METHODOLOGY · Cloud: tested 2026-08-02 on 12 university lectures via OpenRouter, using LectureSync's production note-taking prompt (single pass, low temperature). Notes scored automatically against a curated, hand-built answer key per lecture: faithfulness (fabricated facts), detail (worked examples & specific quantities retained), coverage (key points on unfamiliar subjects), and on-topic discipline. Every model listed above was then re-run 300 times on our hardest math lectures, and every automatic fabrication flag was re-derived by hand against the transcript. Faithfulness limit: 0.5% adjudicated fabrication rate: above it, a model is excluded whatever its other scores. Costs are real OpenRouter token charges projected to a 75-lecture semester (five a week for fifteen weeks); where an endpoint is free-only, we quote the vendor's published paid list rate and mark it ‡. Local: multi-seed study on GGUF models, MIT OCW lectures; full scores on the local model leaderboard.
Last updated: August 2, 2026 · Local: 1,755-run bake-off · Cloud: 31 models, 300 runs per finalist (2026-08-02)