Ask Cat › AI Tool Summary › Otter.ai
US$0.10 an Hour to Transcribe Audio: Microsoft's MAI-Transcribe-2 Promo Ends This Year, and Most People Still Should Not Use It
Article last updated:2026-09-04
Microsoft launched its speech-to-text model MAI-Transcribe-2 on September 3, 2026. The official model page lists it at US$0.10 per hour of audio.
How low is that? In a more familiar unit: roughly US$1.67 per 1,000 minutes. A full eight-hour day of meetings costs under US$1 to transcribe.
There is a caveat, and Microsoft prints it themselves: limited-time.
Last verified: 2026-09-04
1. The price is promotional, and the real one is not public
The official model page states plainly that US$0.10/hour is a limited-time offer. The same page cites the previous generation, MAI-Transcribe-1.5, at US$0.36 per hour for comparison.
Press coverage (VentureBeat, Neowin) adds that the promotional price runs only through the end of 2026, and that Microsoft has not disclosed the post-promotional rate.
The practical consequence: if you are building a cost model for a production system, the number you compute today may not hold in January. Extrapolating from the US$0.36 predecessor, a return to something three times higher is plausible — but that is inference, not published fact, so we do not state it as one.
Evidence levels: the US$0.10/hour figure and the “limited-time” label come from Microsoft’s official model page. “Through the end of 2026, regular price undisclosed” comes from VentureBeat and Neowin; we did not find an explicit end date on the official page.
2. Speed: one hour of audio in about ten seconds
The official page gives the figure as 1hr audio → 10 sec of inference.
Reported comparisons: roughly 10x faster than OpenAI’s GPT-Transcribe, 7x faster than ElevenLabs’ Scribe v2, and 5x faster than Google’s Gemini 3.5 Transcribe.
For an individual, the difference between 10 seconds and 60 seconds is not decisive — you upload the recording, you make coffee, both finish. Speed matters at volume: a contact center processing thousands of calls a day, a podcast network backfilling captions, a compliance team sweeping three years of recordings. Ten-times-faster per item becomes ten-times-cheaper compute and ten-times-shorter queues in batch.
3. Accuracy: two different numbers, and they are not the same test
The official page and the press cite different accuracy figures. Keeping them separate:
- Official model page: on the FLEURS benchmark’s top 25 languages, average word error rate (WER) of 3.4%; it also cites an Artificial Analysis ranking of #2 at 2% WER.
- Press (VentureBeat / Neowin): across 60 FLEURS languages, average WER of 5.2%, ranked first.
These do not contradict each other — wider language coverage raises average error, because low-resource languages are harder. Which number you quote determines how miraculous it sounds. We list both.
What does 5.2% WER feel like? Roughly one wrong word in twenty. The wrong one is usually a name, a technical term or a number — so transcripts still need a human pass before they go anywhere external. No model generation changes that.
4. The features that actually matter in production
Three items on the official feature list make a real difference to workflows:
- Diarization — labels who said what. Meeting notes without it are soup.
- Word-level timestamps — every word carries a time position, which is what editing, captioning and “where did she say that” all depend on.
- Domain biasing — feed it a list of product names, drug names or people, so it stops rendering your company name as something else. This is the biggest practical win in specialist settings.
Microsoft also cites optimization for background noise and imperfect recordings. Language support: 60 languages. Access: Microsoft Foundry and the MAI Playground.
5. Honestly: most people should not touch the API
This is the important part.
US$0.10 per hour is an API price. It assumes you write the calls, handle uploads, store results and build an interface. For a developer that is cheap. For someone who just wants meeting notes, the cost is not money — it is weekends.
Compare with the ready-made tools we have verified:
- Otter.ai: free tier gives 300 transcription minutes per month, 30 minutes maximum per meeting (read from the official pricing page 2026-07-31).
- Fathom: free tier gives unlimited recording, transcription and storage, but advanced AI summaries only for the first 5 meetings each month (official pricing page, 2026-08-28). Note its 38 supported transcription languages do not include Chinese.
- Fireflies: unlimited transcription on free, but capped by 400 minutes of storage, and transcript downloads are paid-only.
The dividing line is simple: dozens of hours a month plus existing engineering capacity → the API pays for itself. Three meetings a week and you want an app to click → a free tier almost certainly covers you.
Related
- Otter.ai free tier and plans, verified
- Fathom free plan: what unlimited actually means
- Fireflies free plan: the 400 minutes do not reset
- Gemini 3.5 Transcribe launch specs
Sources
- Microsoft AI official model page: MAI-Transcribe-2
- VentureBeat: Microsoft AI’s MAI-Transcribe-2 undercuts OpenAI, Google and ElevenLabs on price and speed (2026-09-03)
- Neowin: Microsoft’s MAI-Transcribe-2 model beats OpenAI and Google while costing just US$0.10 per hour (2026-09-03)
- This site’s Otter.ai / Fathom / Fireflies verification logs
Independent reporting, not sponsored. Check Microsoft’s official pricing page before budgeting.
What Amo and Pimi think
For meeting notes needs: try the free plan's 300 minutes/month first. If you're in meetings more than an hour a week: Pro at US$8.33/month with annual billing is the best value. For unlimited team use: Business at US$19.99 with annual billing. But Chinese transcription quality hasn't been actually tested — try before you buy.
Let's take a look at these
- Otter.ai Comprehensive Introduction: Pricing, Features, and Actual Limitations
- Otter.ai Is the free quota enough?
- Otter.ai Alternatives
- Comprehensive Free Quota List for All Tools

