1 / 1

just talk

why voice is the next interface. and how we built one

sanju

sanju

thisux  ·  unitedby.ai

how many of you talked to a machine this week?

siri, chatgpt, gemini, the car, whatever. raise a hand.

and how many typed into a box instead?

yeah. that's most of us. that's the gap.

"

voice is the oldest interface we have.

we just forgot, because screens showed up.

why voice?

not because it's trendy. because it's how we already work.

every person in this room

learned to talk
before they learned
to type.

the timescales are not even close

200k

years of speech

our brains are built for this

150

years of typing

a skill we had to learn

50

years of screens

a blip. a very loud blip.

"

talking is a reflex.

typing is a skill.

stanford, 2016. phones. english and mandarin.

3×

faster to speak than to type

Ruan, Wobbrock, Liou, Ng, Landay · Stanford HCI

same study. english. words per minute.

speaking

153

wpm

typing on a phone

52

wpm

and speech had fewer errors while you were entering it.

speed is the easy argument.

voice also carries
what text kills.

tone. pause. heat. doubt. a laugh you didn't mean to let out.

read this out loud in your head.

i'm fine.

now say it like you mean it. now say it like you don't.
same two words. totally different product.

when two humans talk, the gap between turns is

200ms

that's one syllable.

Stivers et al., PNAS 2009 · 10 languages. mean gap ~208ms.

if your agent takes two seconds to answer,

it doesn't feel slow.

it feels broken.

and a lot of life is not a laptop

driving

hands on the wheel. eyes on the road. voice is the only safe ui.

cooking, warehouse, clinic

gloves on. messy hands. you talk, you don't tap.

kids, parents, tired people

nobody wants to hunt a setting. they just want to say it.

companions

long sessions. late nights. text feels like email. voice feels like someone.

accessibility is not a feature.

for a lot of people, voice is the only interface that works.

"

the best interface is the one you already know.

you already know how to talk.

how we actually
use products now

the computer keeps getting closer to the human

every era, the computer moved one step closer

to how humans
already work.

five interfaces. one direction.

cli

you learn the machine's language

gui

you point at pictures

touch

you poke the thing itself

chat

you type like you'd text a friend

voice

you just talk

2022

we all started chatting with boxes.

that was the first time software felt like a person, not a form.

then the boxes started talking back

chatgpt voice.
gemini live.
gpt live.

alexa, but it can actually think.

you don't open an app.

you ask.

this is the quiet shift

products are becoming conversations.

not a screen with buttons. a turn. then another turn.

the old product vs the new one

from

screens

to sessions

from

buttons

to turns

from

pages

to dialogue

us voice assistant users, end of 2024

145M

heading toward 168 million by 2029

eMarketer voice assistant forecast

conversational ai market

~$15B

in 2025. around $18B in 2026.

Fortune Business Insights. other firms land nearby, not identical. treat it as a range, not a gospel.

and people are already handing work to agents

search

"what's the best one for me" happens before they hit your site.

compare

the agent reads the fine print. the human hears the answer.

maybe buy

trust is the bottleneck. talk is how you earn it.

Visa, Accenture, L.E.K. 2025 to 2026. lots of people already use AI to shop. fewer want it to spend money yet.

"

the product is the conversation.

everything else is plumbing.

what a voice agent
actually is

spoiler: not magic. a loop.

it is not one model.

it is a bunch of systems that have to finish inside that 200ms feeling.

the loop. say it with me.

mic
→
stt
→
llm
→
tools
→
tts
→
speaker

speech in. words. thought. maybe an action. speech out.

that's the demo. this is the product.

transport

webrtc, websocket, sip, the phone network

session

who is this, where are we, did we drop

vad

are they talking, or is that a truck

barge-in

they talked over you. now what.

memory

don't ask their name twice

logs

when it fails at 2am, you need a story

the naive way

wait for the full sentence.
wait for the full answer.
wait for the full audio.

2 to 4 seconds.

conversation dies. they hang up. they type instead.

"

latency is the product.

the model is just one of the delays.

the parts that
actually hurt

this is where most voice apps fall over

five problems. if you miss one, it feels fake.

01

latency. stream everything, or don't bother

02

turn-taking. when are they actually done talking

03

interruptions. they will talk over you. humans do.

04

tools. don't go silent while you "check"

05

lock-in. your stt vendor will change. they always do.

stream everything.

bad

full audio → full transcript → full reply → full speech. four waiting rooms.

good

partials in. tokens out. first sentence hits the speaker while the rest is still thinking.

when are they done?

silence is not the same as "your turn".

a pause

they're thinking. cut in and you look rude.

the end

wait too long and you look asleep.

this is the whole game

"your meeting starts..."

"no, tomorrow."

if the agent finishes the sentence, you built a radio. not a conversation.

two ways to hear them

half duplex

walkie-talkie

you talk. then they talk. if they jump in, you stop, go back to listening, start over.

full duplex

a real call

both sides stay live. they talk over you. you hear it. you change what you were going to say.

adapt.
don't restart.

keep what you already said. fold in the correction. keep going.

don't go silent.

"let me check…" and then two seconds of nothing is how you lose them.

hold first speech

if a tool is coming, don't start a sentence you can't finish.

then stream

once the tool is back, talk as the words arrive.

your vendor will change.

stt, llm, tts. prices move. quality moves. one of them will annoy you.

if swapping groq for openai means rewriting the app, you don't have an architecture. you have a demo.

"

if it feels like a phone tree, you lost.

people hang up on phone trees.

pipecat

this is how most serious voice agents get built today

what it is, in plain words

an open source python framework for real-time voice agents.

built by Daily. bsd licensed. you bring the models. they give you the pipe.

the idea that made it click

everything is a frame.

audio is a frame. a transcript is a frame. a tool result is a frame. they flow.

a typical pipecat pipeline

transport.in

mic, webrtc, the phone

stt

audio → words

context + llm

words → thought, maybe a tool

tts

thought → audio

transport.out

speaker, the other end of the call

three kinds of frames. this is the clever bit.

system

high priority. start, stop, interrupt. these jump the queue. they don't get cancelled.

data

the actual stuff. audio, text, images. if the user jumps in, these get dropped.

control

config and lifecycle. "start talking". "that's the end of the sentence".

why people actually use it

composable

swap deepgram for openai. the rest of the pipe stays.

real-time first

streaming is the default, not a later patch.

transports included

daily, websocket, twilio. the hard audio path is not your problem.

a real community

this is not a weekend repo. daily ships it, people run it in prod.

the catch, for a lot of us

python.

our apps are typescript. our edge is hono and workers. our team does not want a second runtime just to talk.

"

we needed this, in the language we ship in.

so we built it.

thisux voice

a typescript sdk for real-time voice agents

not another voice framework. an infrastructure sdk.

same category as the tools you already trust

email

resend

files

uploadthing

storage

unstorage

auth

better auth

voice

thisux voice

the happy path is one call

createVoice({
stt, llm, tts })

providers as config. tools as plugins. events for everything else.

that's it. that's the agent.

const voice = createVoice({
  transport: webrtc(),
  stt: openai(),
  llm: groq(),
  tts: cartesia(),
});

voice.tool({
  name: "createTask",
  async execute({ title }) {
    return { id: crypto.randomUUID(), title };
  },
});

voice.on("transcript.final", ({ text }) => {
  console.log("user:", text);
});

await voice.connect();

what's under that call

your app

tools, business logic

createVoice

public api, middleware, events

session

lifecycle, reconnect, state machine

pipeline

audio → stt → llm → tools → tts

providers

swap these. don't rewrite the rest.

providers are config. not architecture.

stt

openai

deepgram later. same interface.

llm

groq, openai

change one line. keep the tools.

tts

cartesia, elevenlabs

pick the voice. not the stack.

same agent. four ways in.

webrtc

browsers. the default. low latency.

websocket

when you just need frames, not ice.

sip

real phones. the old network, still huge.

twilio

media streams in, agent out. support teams live here.

everything emits. you subscribe.

transcript.partial

they're still talking

transcript.final

that's the sentence

llm.started

thinking begins

tool.called

it reached for your api

tts.started

audio is leaving

duplex.overlap

they talked over us. we adapt.

the session is a state machine. not a pile of flags.

idle → connecting → connected → listening
listening → thinking → speaking → listening

speaking → thinking → speaking     // adapt. no restart.
speaking → interrupted → listening // hard stop, if you want it

* → reconnecting → connected | failed
* → closed

the techniques, all in one place

provider interfaces

stt / llm / tts / transport. new vendor = new package, same contract.

full duplex adapt

mic stays open while we speak. overlap continues the turn.

energy barge-in

loud inbound pcm stops leftover tts. not full aec. a courtesy stop.

sentence-flush tts

llm tokens → first period → audio. don't wait for the essay.

abort everywhere

tts.abort(). tools get AbortSignal. leftover work dies with the turn.

middleware

voice.use(logger()). voice.use(metrics()). better-auth energy.

edge-safe core

no node-only apis in core. hono and workers are first-class.

offline fakes

tests run with no api keys. the suite does not need groq to be up.

duplex, the default

createVoice({
  duplex: {
    listenWhileSpeaking: true,
    onOverlap: "adapt",     // or "interrupt"
    overlapTimeoutMs: 4000,
  },
  bargeIn: {
    energyThreshold: 0.025,
    graceMs: 250,           // ignore echo right after tts
  },
});

they talk over you → leftover audio stops → we keep context → we answer the new thing.

streamed tts. this is how you beat 2 seconds.

flush on the period.

llm is still writing sentence two. sentence one is already in their ear.

default

ttsStreaming: true

safety

if tools might fire, we hold the first pass so we don't say "let me check" and then vanish.

what we are not building. on purpose.

not a model lab

we don't train speech models. we don't own gpus.

not a phone company

we don't own a telephony network. we plug into one.

not livekit's job

we don't ask you to wire 12 primitives for the happy path.

not a python dsl

pipecat is great. this is for the typescript side of the house.

"

swap a provider.

don't rewrite the app.

take these with you

thisux voice

the sdk this talk is about

github.com/thisuxhq/voice

pipecat

the python pipeline we learned from

docs.pipecat.ai

200ms turn gap

why latency is the whole product

pnas.org · stivers 2009

3× faster

speech vs typing on phones

stanford hci · ruan 2016

if you remember one thing

people will talk to your product the way they talk to each other.

your job is to not make that feel weird.

Q&A

ask anything. talk. that's the point.

sanju

sanju

thisux.com  ·  unitedby.ai  ·  sanju.sh

thisux voice  ·  september 2026