Access powerful AI models to transcribe and understand speech via a simple API. Try our no-code playground for free 👉 https://nitter.cf/t.co/YPCK9mqDG6

Joined October 2017
Include

Only show posts containing:

Exclude

Hide posts containing:

Time range
-
Minimum likes
We can’t wait to see you all for the SF Voice AI meetup today! Casual AI chats and you take a voice app home! 3 hours to go, sign up below
4
7
1,483
SF Voice AI Meetup today! Excited to have you all here! Sign up link below:
2
1
22
1,880
AssemblyAI Universal 3.5 Pro topped our speech-to-text benchmark on both accuracy and latency. We ran 15 models through the same audio and pipeline. On 1,000 read-aloud clips from Pipecat FLEURS, Universal 3.5 Pro had the lowest word error rate at 1.93%. It also had the fastest median time to first text at 489ms, so it started returning a transcript sooner than any other model we tested. Congrats to the @AssemblyAI team. Full results: benchmarks.cekura.ai/stt
2
4
21
5,945
8 more days left: we can’t wait to see what you all build!
Close your laptop. The agent keeps taking calls. New tutorial: a voice agent on AssemblyAI's Voice Agent API with HTTP tools. No dispatcher, no WebSocket in your backend. AssemblyAI calls your API itself, mid-call. lablab.ai/ai-tutorials/assem….
4
4
1,567
Can’t wait to see what amazing voice apps you build in 15 mins at the workshop! See you all in San Francisco: luma.com/xwnkujzr
The fastest way to get text into a computer is to talk.If your product has a text field, it has somewhere for dictation to go. Thursday in SF, @iHarnoorSingh shows you how to put it there in about 30 lines of Python. Save your seat: luma.com/xwnkujzr
1
3
1,260
AssemblyAI retweeted
Universal-3.5 Pro from @AssemblyAI is live on OpenRouter, 50% off through September 29! Speech-to-text in 19 languages in a single synchronous call. Ranks first for accuracy across independent speech-to-text benchmarks. Send an audio clip of up to two minutes and quickly get the full transcript back with word-level timestamps and confidence scores. Steer transcription toward your domain by supplying keyterms and contextual prompting. Try it: openrouter.ai/assemblyai/uni…
7
5
1
131
22,353
The fastest way to get text into a computer is to talk.If your product has a text field, it has somewhere for dictation to go. Thursday in SF, @iHarnoorSingh shows you how to put it there in about 30 lines of Python. Save your seat: luma.com/xwnkujzr
2
1
4
10,503
SF, check your mailboxes 👀 Invites to our SF Tech Week Voice Agent workshop went out this week. Not sure "vintage" feels quite right, but we can't wait to see what you code into your voice-agent sidekick to clip onto your backpack and relive your childhood dreams. Didn't get one? Request an invite 👇 luma.com/w9e4qgol
4
3
9
10,559
BTW, we intentionally spelled Minoxidil as “Minoxidril”; the model spells it exactly as provided in the keyterms.
2
234
Blurt: push‑to‑talk dictation on the AssemblyAI Dictation API The Dictation API converts a spoken clip into finished text. Filler words and false starts are removed, and the output can be formatted as notes, a commit message, or a reply to a customer. Built on Universal‑3.5 Pro, it supports 19 languages, processes short clips in under a second, and costs $0.62/hr. Link in comments below
7
2
12
1,717
Fingers can't type as fast as your brain moves? That's why the little mic icon is showing up everywhere: notes apps, email, CRMs, coding tools, patient charts. People talk, the app writes it down. On Sept 24 @iHarnoorSingh is showing SF how to build that feature with our Dictation API. Save your seat → luma.com/xwnkujzr?utm_source…
3
2
9
11,704
Dear developers, Forms are broken. Questions across multiple pages and endless fields. A slow, painful experience. What if we could answer everything… with a single voice prompt? With dictation, users can input text up to 3× faster. That's why, we’ve launched our Dictation API: assemblyai.com/dashboard/pla…
1
3
1
6
2,488
"um so like can you send me the uh the Q3 numbers" is what the mic heard. "Can you send me the Q3 numbers?" is what the person meant. @iHarnoorSingh walks through the Dictation API playground showing the gap close. Add your key terms, hold to record, read what comes back. Two things worth watching: The key terms field. Type in the vocabulary your users actually say (drug names, SKUs, your own product names) and those words come back spelled right instead of phonetically close. That's the part most dictation projects die on. The turnaround. Typical short clips come back in under a second, on Universal-3.5 Pro, across 19 languages.. Bring your own audio and throw the words your users get wrong at it 👇 assemblyai.com/dashboard/pla…
3
9
7,975
Sub-200ms latency is hard to picture. Easier to watch. Here's @iHarnoorSingh using the Dictation API. Hotkey down, normal speaking voice, and the text is ready before he looks up from the keyboard. Try the Dictation API yourself: assemblyai.com/blog/dictatio… What would you put dictation into first?
3
2
1
6
17,115
Which of your apps do you wish you could just talk to? Hold a hotkey, talk, let go. The text is there. The Dictation API is live today. It's built for that loop, and for apps where users run it a few hundred times a day. • 134ms p50, so text lands about as fast as you'd type it • Punctuation and casing already applied when it comes back • One endpoint to integrate, so you can ship dictation this sprint Start building: assemblyai.com/blog/dictatio…
6
15,625
Dictation is easy to prototype and hard to ship. Making the speech-to-text → cleanup loop feel instant, at a price that survives production volume, is the actual work. Our forward-deployed engineers walk through it live, Wed Sept 9. Plus, get $20 in credits to build it yourself. Register: assemblyai.zoom.us/webinar/r…
3
1
1
10
24,783
Replying to @dan_aai
@dan_aai builds voice agents at AssemblyAI, which means he thinks about how they're actually put together more than most people do. His framing: a voice agent isn't one model—it's two roles. 🔹 The Responder holds the call almost the entire time—voice activity detection, streaming transcription on Universal-3.5 Pro Realtime, turn detection, TTS, barge-in. 🔹 The Thinker decides what to say. It comes into existence when turn detection fires, and it's gone the moment the last token is written. The Responder is the hard, latency-bound part. Nobody picks your product because your barge-in is 40ms tighter—but they'll leave if the agent talks over them. The Thinker is your domain logic, your data, your escalation rules. That's the product. And in our Voice Agent API it's one config field pointing at a URL you own. Dan's demo repo proves it the least flattering way possible: two of its three "models" are a handful of regexes and a person clicking buttons. The Responder can't tell the difference and doesn't try. Read the deep dive: assemblyai.com/blog/thinker-… Clone the demo and take a live call yourself: github.com/dan-ince-aai/voic…
9
2
1
15
2,085
If you're building dictation: our forward-deployed engineering team is hosting a workshop on building and optimizing dictation features on Wednesday, September 9—live examples, the full pipeline, Q&A throughout. Register: assemblyai.zoom.us/webinar/r…
2
2
7
1,736
Qwen3.5 4B is now on AssemblyAI's LLM Gateway—optimized and hosted by our team for the fast rewrite tasks at the center of voice products: dictation cleanup, transcript rewrite, live formatting. On voice rewrite tasks it averages 612ms—1.9× faster than GPT-4.1, at 94% lower cost per hour of audio on average. Not every task needs a frontier model. Dictation cleanup needs speed, at a cost that holds up at production volume. That's the job this model was purpose-built for. Try it on your own voice tasks today—Qwen3.5 4B is available to test with free credits on your AssemblyAI account, no credit card required. Read the launch post: assemblyai.com/blog/qwen-4b-…
6
1
1
10
11,585