why voice is the next interface. and how we built one
sanju
thisux · unitedby.ai
how many of you talked to a machine this week?
siri, chatgpt, gemini, the car, whatever. raise a hand.
and how many typed into a box instead?
yeah. that's most of us. that's the gap.
"
voice is the oldest interface we have.
we just forgot, because screens showed up.
not because it's trendy. because it's how we already work.
every person in this room
learned to talk
before they learned
to type.
the timescales are not even close
200k
years of speech
our brains are built for this
150
years of typing
a skill we had to learn
50
years of screens
a blip. a very loud blip.
"
talking is a reflex.
typing is a skill.
stanford, 2016. phones. english and mandarin.
3×
faster to speak than to type
Ruan, Wobbrock, Liou, Ng, Landay · Stanford HCI
same study. english. words per minute.
speaking
153
wpm
typing on a phone
52
wpm
and speech had fewer errors while you were entering it.
speed is the easy argument.
voice also carries
what text kills.
tone. pause. heat. doubt. a laugh you didn't mean to let out.
read this out loud in your head.
i'm fine.
now say it like you mean it. now say it like you don't.
same two words. totally different product.
when two humans talk, the gap between turns is
200ms
that's one syllable.
Stivers et al., PNAS 2009 · 10 languages. mean gap ~208ms.
if your agent takes two seconds to answer,
it doesn't feel slow.
it feels broken.
and a lot of life is not a laptop
driving
hands on the wheel. eyes on the road. voice is the only safe ui.
cooking, warehouse, clinic
gloves on. messy hands. you talk, you don't tap.
kids, parents, tired people
nobody wants to hunt a setting. they just want to say it.
companions
long sessions. late nights. text feels like email. voice feels like someone.
accessibility is not a feature.
for a lot of people, voice is the only interface that works.
"
the best interface is the one you already know.
you already know how to talk.
the computer keeps getting closer to the human
every era, the computer moved one step closer
to how humans
already work.
five interfaces. one direction.
cli
you learn the machine's language
gui
you point at pictures
touch
you poke the thing itself
chat
you type like you'd text a friend
voice
you just talk
2022
we all started chatting with boxes.
that was the first time software felt like a person, not a form.
then the boxes started talking back
chatgpt voice.
gemini live.
gpt live.
alexa, but it can actually think.
you don't open an app.
you ask.
this is the quiet shift
products are becoming conversations.
not a screen with buttons. a turn. then another turn.
the old product vs the new one
from
screens
to sessions
from
buttons
to turns
from
pages
to dialogue
us voice assistant users, end of 2024
145M
heading toward 168 million by 2029
eMarketer voice assistant forecast
conversational ai market
~$15B
in 2025. around $18B in 2026.
Fortune Business Insights. other firms land nearby, not identical. treat it as a range, not a gospel.
and people are already handing work to agents
search
"what's the best one for me" happens before they hit your site.
compare
the agent reads the fine print. the human hears the answer.
maybe buy
trust is the bottleneck. talk is how you earn it.
Visa, Accenture, L.E.K. 2025 to 2026. lots of people already use AI to shop. fewer want it to spend money yet.
"
the product is the conversation.
everything else is plumbing.
spoiler: not magic. a loop.
it is not one model.
it is a bunch of systems that have to finish inside that 200ms feeling.
the loop. say it with me.
speech in. words. thought. maybe an action. speech out.
that's the demo. this is the product.
transport
webrtc, websocket, sip, the phone network
session
who is this, where are we, did we drop
vad
are they talking, or is that a truck
barge-in
they talked over you. now what.
memory
don't ask their name twice
logs
when it fails at 2am, you need a story
the naive way
wait for the full sentence.
wait for the full answer.
wait for the full audio.
2 to 4 seconds.
conversation dies. they hang up. they type instead.
"
latency is the product.
the model is just one of the delays.
this is where most voice apps fall over
five problems. if you miss one, it feels fake.
01
latency. stream everything, or don't bother
02
turn-taking. when are they actually done talking
03
interruptions. they will talk over you. humans do.
04
tools. don't go silent while you "check"
05
lock-in. your stt vendor will change. they always do.
bad
full audio → full transcript → full reply → full speech. four waiting rooms.
good
partials in. tokens out. first sentence hits the speaker while the rest is still thinking.
silence is not the same as "your turn".
a pause
they're thinking. cut in and you look rude.
the end
wait too long and you look asleep.
this is the whole game
"your meeting starts..."
"no, tomorrow."
if the agent finishes the sentence, you built a radio. not a conversation.
two ways to hear them
half duplex
walkie-talkie
you talk. then they talk. if they jump in, you stop, go back to listening, start over.
full duplex
a real call
both sides stay live. they talk over you. you hear it. you change what you were going to say.
adapt.
don't restart.
keep what you already said. fold in the correction. keep going.
"let me check…" and then two seconds of nothing is how you lose them.
hold first speech
if a tool is coming, don't start a sentence you can't finish.
then stream
once the tool is back, talk as the words arrive.
stt, llm, tts. prices move. quality moves. one of them will annoy you.
if swapping groq for openai means rewriting the app, you don't have an architecture. you have a demo.
"
if it feels like a phone tree, you lost.
people hang up on phone trees.
this is how most serious voice agents get built today
what it is, in plain words
built by Daily. bsd licensed. you bring the models. they give you the pipe.
the idea that made it click
everything is a frame.
audio is a frame. a transcript is a frame. a tool result is a frame. they flow.
a typical pipecat pipeline
transport.in
mic, webrtc, the phone
stt
audio → words
context + llm
words → thought, maybe a tool
tts
thought → audio
transport.out
speaker, the other end of the call
three kinds of frames. this is the clever bit.
system
high priority. start, stop, interrupt. these jump the queue. they don't get cancelled.
data
the actual stuff. audio, text, images. if the user jumps in, these get dropped.
control
config and lifecycle. "start talking". "that's the end of the sentence".
why people actually use it
composable
swap deepgram for openai. the rest of the pipe stays.
real-time first
streaming is the default, not a later patch.
transports included
daily, websocket, twilio. the hard audio path is not your problem.
a real community
this is not a weekend repo. daily ships it, people run it in prod.
the catch, for a lot of us
python.
our apps are typescript. our edge is hono and workers. our team does not want a second runtime just to talk.
"
we needed this, in the language we ship in.
so we built it.
a typescript sdk for real-time voice agents
not another voice framework. an infrastructure sdk.
same category as the tools you already trust
resend
files
uploadthing
storage
unstorage
auth
better auth
voice
thisux voice
the happy path is one call
createVoice({
stt, llm, tts })
providers as config. tools as plugins. events for everything else.
that's it. that's the agent.
const voice = createVoice({ transport: webrtc(), stt: openai(), llm: groq(), tts: cartesia(), }); voice.tool({ name: "createTask", async execute({ title }) { return { id: crypto.randomUUID(), title }; }, }); voice.on("transcript.final", ({ text }) => { console.log("user:", text); }); await voice.connect();
what's under that call
your app
tools, business logic
createVoice
public api, middleware, events
session
lifecycle, reconnect, state machine
pipeline
audio → stt → llm → tools → tts
providers
swap these. don't rewrite the rest.
providers are config. not architecture.
stt
openai
deepgram later. same interface.
llm
groq, openai
change one line. keep the tools.
tts
cartesia, elevenlabs
pick the voice. not the stack.
same agent. four ways in.
webrtc
browsers. the default. low latency.
websocket
when you just need frames, not ice.
sip
real phones. the old network, still huge.
twilio
media streams in, agent out. support teams live here.
everything emits. you subscribe.
transcript.partial
they're still talking
transcript.final
that's the sentence
llm.started
thinking begins
tool.called
it reached for your api
tts.started
audio is leaving
duplex.overlap
they talked over us. we adapt.
the session is a state machine. not a pile of flags.
idle → connecting → connected → listening listening → thinking → speaking → listening speaking → thinking → speaking // adapt. no restart. speaking → interrupted → listening // hard stop, if you want it * → reconnecting → connected | failed * → closed
the techniques, all in one place
provider interfaces
stt / llm / tts / transport. new vendor = new package, same contract.
full duplex adapt
mic stays open while we speak. overlap continues the turn.
energy barge-in
loud inbound pcm stops leftover tts. not full aec. a courtesy stop.
sentence-flush tts
llm tokens → first period → audio. don't wait for the essay.
abort everywhere
tts.abort(). tools get AbortSignal. leftover work dies with the turn.
middleware
voice.use(logger()). voice.use(metrics()). better-auth energy.
edge-safe core
no node-only apis in core. hono and workers are first-class.
offline fakes
tests run with no api keys. the suite does not need groq to be up.
duplex, the default
createVoice({
duplex: {
listenWhileSpeaking: true,
onOverlap: "adapt", // or "interrupt"
overlapTimeoutMs: 4000,
},
bargeIn: {
energyThreshold: 0.025,
graceMs: 250, // ignore echo right after tts
},
});
they talk over you → leftover audio stops → we keep context → we answer the new thing.
streamed tts. this is how you beat 2 seconds.
flush on the period.
llm is still writing sentence two. sentence one is already in their ear.
default
ttsStreaming: true
safety
if tools might fire, we hold the first pass so we don't say "let me check" and then vanish.
what we are not building. on purpose.
not a model lab
we don't train speech models. we don't own gpus.
not a phone company
we don't own a telephony network. we plug into one.
not livekit's job
we don't ask you to wire 12 primitives for the happy path.
not a python dsl
pipecat is great. this is for the typescript side of the house.
"
swap a provider.
don't rewrite the app.
take these with you
thisux voice
the sdk this talk is about
github.com/thisuxhq/voice
pipecat
the python pipeline we learned from
docs.pipecat.ai
200ms turn gap
why latency is the whole product
pnas.org · stivers 2009
3× faster
speech vs typing on phones
stanford hci · ruan 2016
if you remember one thing
people will talk to your product the way they talk to each other.
your job is to not make that feel weird.
ask anything. talk. that's the point.
sanju
thisux.com · unitedby.ai · sanju.sh
thisux voice · september 2026