JEV Voice Agent
A spoken-conversation agent whose hard problem is not speech β it is turn-taking. Two independent JEV judgements decide when there is enough to answer, and when the person has actually finished.
Speech recognition and speech synthesis are bought from an API. They are solved, and they are not what this build is about. The hard problem in a spoken conversation is turn-taking: knowing when the person has finished, and knowing when it is your turn to speak.
Those are two different questions, and the central claim of this build is that they have to be asked separately β by two independent judges, looking at two different pieces of text.
ββββββββββββββββββ the ear βββββββββββββββββββ
mic ββΆ STT ββΆ U partials overwrite, finals append
β
β ticker, 1 Hz
βΌ
JEV#1 "is there enough content to answer yet?"
β ββ false ββΆ wait
β true
βΌ
ββββββββββββββββ
β reply list β βββ more cuts pile up while the LLM is busy
ββββββββββββββββ
β
β AND the person has been quiet long enough
βΌ
LLM one turn β every pending cut, joined with spaces
β
β text streams out in chunks
βΌ
GateSink holds the first chunk, and asks:
β
JEV#2 "has this person finished their turn?"
β ββ false ββΆ keep holding
β true
βΌ
MicGate mutes the mic while we are speaking
βΌ
SpeakSink ββΆ TTS ββΆ speaker
βββββββββββββββββ the mouth ββββββββββββββββββ
Why Turn-Taking Is the Whole Problem
A naive voice agent waits for the speech-to-text provider to announce "end of utterance", then replies. That fails in both directions, constantly:
- It interrupts. You say "I have three questions, the first one is about indexes." That is a grammatically complete sentence, so the provider closes the utterance β and the agent starts answering question one while you are drawing breath for question two.
- It stalls. You say "I wanted to ask" and pause to think. Nothing in the transcript closes, so nothing happens, and the silence stretches until you repeat yourself.
Neither failure raises an error. Both just make the conversation feel wrong. So the system needs an actual judgement β and it needs two of them, because the same text gets opposite answers depending on which question you ask:
| JEV#1 β the ear | JEV#2 β the mouth | |
|---|---|---|
| Question | Is there enough content to answer? | Has this person finished their turn? |
| What it reads | the live transcript since the last submission | the whole chunk just submitted, plus what the silence has done since |
| If true | submit into the reply list, start the LLM | open the gate, start speaking |
| "I have three questions, the first is about indexes." | true β that sentence is answerable | false β he obviously has two more |
| "uhβ¦ soβ¦ theβ¦" | false β no content | false β still hunting for a word |
Those two rows are the design. One judge cannot answer both questions, and collapsing them produces exactly the interrupting agent described above.
JEV β A Judgement, Not a Prompt
JEV is a judgement call made by a model through a decisions API: it takes a state string and a question with explicit criteria, and returns a calibrated scalar plus a verdict. It is not a chat completion, and the build does not treat it as one. Three consequences shape the code:
1. The question file is the policy
Both questions live in version control as authoritative files, and each contains exactly one fenced JSON block. The client reads only that block β everything around it is reasoning written for humans and never reaches the model. Editing the block is editing behaviour, so it is reviewed like code.
2. The examples are hand-written, never copied from the corpus
If you paste real recorded utterances into the question and then measure accuracy against those same recordings, you have measured recitation. The entire corpus stays held out as the test set, and the benchmark enforces the separation.
3. A verdict that cannot be read is an error, never a false
If the response comes back without a readable scalar, the client raises. Degrading to "not complete" instead would produce a system that never opens its mouth while reporting no error anywhere β the worst failure mode available.
Two further details that cost real debugging to learn:
- Thresholds are scanned, not chosen. JEV#1's threshold of 0.65 sits in the middle of a plateau found by sweeping a labelled corpus. JEV#2's threshold of 0.6 has never been scanned β it is a guess, and it is annotated as a guess in the config and listed under Known Limits rather than quietly shipped as a number.
- A late verdict is discarded, never acted on, and never cancelled. Requests go out roughly once a second and can return out of order; by the time a slow one comes back, the transcript it looked at is stale, and acting on it would submit a fragment as though it were complete. The in-flight request is not cancelled, because cancelling poisons the connection pool β measured, the next call goes from 352β421 ms to 1119β1425 ms.
The Reply List
The reply list is the spine of the build. It exists because speech does not arrive in neat turns, and because a crashed process must not silently lose what someone said. Every piece of state is a file in the session directory, so the conversation survives the process dying.
One utterance, three cuts, one turn
You say something, pause, add to it, pause, add again. JEV#1 passes three times, so three cuts enter the reply list. They do not become three LLM calls:
cut 1 ββ
cut 2 ββΌβββΆ one drain βββΆ "cut one cut two cut three" βββΆ one LLM turn
cut 3 ββ
The cuts are joined with spaces, not newlines β newlines nudge the model into answering each fragment separately. And when cuts arrive while the LLM is already working, they simply accumulate and are eaten in one go by the next turn. That behaviour is what the whole structure exists to produce.
The five pieces
| Piece | Job | The failure it prevents |
|---|---|---|
| Reservoir | the reply list itself β cuts land here, a turn takes the whole batch | the batch is taken by an atomic rename, so a cut written at the same instant cannot be lost |
| UserBuffer | the live transcript β a partial overwrites the current segment, a final freezes it and starts a new one | empty partials and empty finals really do arrive from the provider; unguarded, the first wipes a half-spoken sentence and the second grows the file forever while reading perfectly normally |
| SilenceCounter | how long the person has actually been quiet | it reads the audio timeline the provider stamps on each event, not arrival time, and clamps the correction so clock drift can only under-report |
| SpokenLedger | what was really spoken, versus what was merely generated | reconciliation only ever truncates β over-reporting would write words into the transcript that never left the speaker, producing a lying record |
| VoiceDriver | one pump call = exactly one turn, start to finish | β |
A turn has exactly one ending
Earlier versions had three: finish, get interrupted, get invalidated. All of them are gone. Once a turn starts, it finishes and it is written down. If you speak while the machine is speaking, your words go into the reply list and wait β they do not cut it off.
This is deliberate, and it is defended at the source level, because the deleted machinery is
easy to reinstate by accident: restoring one interrupt event turns no behavioural assertion
red, since nobody sets it, but the next person to read the code will wire the trigger back
up. So four assertions parse the source with Python's ast module and fail if the
names come back.
Three places deliberately throw away what a person said
All three leave a trace, because the one thing worse than dropping a sentence is dropping it silently:
- Three strikes. The transcript is judged insufficient three ticks running with no growth. The "no growth" clause is the entire safeguard against cutting off someone who is still talking.
- The delivered watermark. JEV#1 passes on a partial, the buffer is cleared, and then that block's final arrives and writes the same sentence back into the empty buffer. Without a watermark, the same sentence gets answered twice β it happened in two turns out of six on a real session. The watermark cuts only a verbatim prefix, never a character of new text, and declines to cut at all when coverage drops below threshold: better to repeat a few words than to bite the head off a sentence.
- The hallucination gate. In one session, three seconds after the mic opened with nobody speaking, the provider emitted a confident sentence over a 0.62-second audio window. It drove a full LLM turn, a full TTS turn, and was written permanently into memory. The test is speech rate (11.31 units/sec, against a real person's fastest 4.31), not content, so the gate does not drift with the topic β and it defaults to letting things through when duration is missing, because losing one hallucination is cheaper than swallowing whole batches of real speech.
The Mouth
The reply does not go straight to the speaker. The first chunk is held at a gate while JEV#2 decides whether you have finished, across three bands:
| Real silence | What happens | Cost |
|---|---|---|
| under 1.0 s | hold β do not even ask | free |
| between the bounds | ask JEV#2, rate-limited to once per extra second of silence | one request |
| 6.0 s or more | open unconditionally β do not ask | free |
Checking the gate is free; asking JEV#2 costs money. Without the rate limit, dropping the mouth's tick interval from 1.0 s to 0.25 s would have quadrupled the bill while running perfectly and raising nothing. Once the gate opens it forwards transparently β only the first chunk of a turn ever waits.
MicGate then mutes the microphone for exactly as long as the machine is speaking, plus a short tail for room reverberation. Without it, the shape is: speech out of the speaker β into the mic β transcribed β drives another turn β a full LLM and a full TTS every lap, spending real money β and since nothing interrupts anything, it sounds like a normal conversation while it happens.
It sits inside the gate and outside the speaking sink.
Outside the gate it would mute the mic while the gate is still waiting, destroying the very
silence measurement the gate depends on. It releases in a finally, because one
exception leaving the mic muted means permanently deaf, with nothing on screen to say so.
MicGate is not acoustic echo cancellation. It stops the machine hearing itself. It does not stop it hearing another person, or a second device playing audio. The script asks you to confirm your setup before opening the mic, every time.
β from the repository README
Memory
Conversation is appended to a per-peer thread file by the script, deterministically β the model generates text, it never writes the record. The file survives the process exiting, which also means the last run's conversation is context for the next one.
A retention window of recent blocks is injected verbatim every turn, sized by a token budget rather than a block count, with a hard floor underneath it: a single long monologue makes a pure budget compute zero blocks, and the agent then cannot remember the sentence you just said. Older blocks are compressed into a summary on a worker thread β running that blocking call inline jams the event loop, and for those seconds nobody consumes mic frames. The file is archived before every overwrite, because compression is lossy and that file is the only store of the verbatim words.
The Gate β 151 Offline Assertions
Everything above runs under a regression gate with 151 assertions: no LLM, no network, no microphone, roughly a tenth of a second, and it writes nothing outside a sandbox. It is the main reason the repository is worth reading, because nearly every assertion is a regression for a specific failure that ran, produced no error, and sounded fine.
Each assertion carries its own reasoning inline, so the file doubles as the design record β why barge-in was removed, why the silence reading is measured in seconds rather than ticks, why a stub's duration has to be physically plausible.
The structural technique worth stealing: where an assertion must check
that a deleted mechanism has not returned, it parses the source with ast
rather than grepping. The prose explaining a deletion necessarily names the deleted
thing β so a grep would match its own documentation and stay permanently red.
Known Limits
Stated plainly, because none of these announce themselves at runtime.
- JEV#2's threshold is a guess. JEV#1's was scanned across a labelled corpus; JEV#2's never was. The benchmark that would fix it has not been built.
- The gate's lower bound sits below the mid-sentence peak. The silence reading naturally climbs to 1.6β2.2 s while someone is still speaking, so at 1.0 s the gate often asks JEV#2 mid-sentence β 36 of 38 measured within-sentence gaps. A knowing trade-off: raise it toward 2.0 to tighten, and change nothing else.
- Every calibrated number was measured on recordings of one person speaking Chinese. The speech-rate ceiling, the pipeline lag, the thresholds, the retention budget. The shapes of the defects are language-independent and several were re-observed verbatim in English, but no number has been re-validated on English speech. Treat them as starting points and re-scan on your own corpus.
- The corpus ships empty. The original recordings are of a specific person and are not distributed. Record and label your own and the benchmark becomes runnable; six corpus-replay assertions were removed from the gate rather than left permanently skipped.
- No automatic STT reconnect, and the TTS session permit expires after 600 s untreated β the symptom is that it quietly stops speaking while every turn still reports success.
Get It
GitHub: IntelligenismCommercialDevelopment-LLC/jev_voice_agent β
Requires Python 3.12 or newer, a microphone, and headphones. Run the offline gate first β it is free and takes about a tenth of a second.
Derived from the AI Operability Edition of the Intelligenism Agent Framework β specialized into a spoken-conversation agent built around two independent turn-taking judgements.