JEV Voice Agent

Speech recognition and speech synthesis are bought from an API. They are solved, and they are not what this build is about. The hard problem in a spoken conversation is turn-taking: knowing when the person has finished, and knowing when it is your turn to speak.

Those are two different questions, and the central claim of this build is that they have to be asked separately β€” by two independent judges, looking at two different pieces of text.

          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€ the ear ──────────────────┐
 mic ─▢ STT ─▢ U      partials overwrite, finals append
          β”‚
          β”‚  ticker, 1 Hz
          β–Ό
       JEV#1   "is there enough content to answer yet?"
          β”‚                        └─ false ─▢ wait
          β”‚ true
          β–Ό
   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
   β”‚  reply list  β”‚ ◀── more cuts pile up while the LLM is busy
   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
          β”‚
          β”‚  AND the person has been quiet long enough
          β–Ό
         LLM     one turn β€” every pending cut, joined with spaces
          β”‚
          β”‚ text streams out in chunks
          β–Ό
      GateSink   holds the first chunk, and asks:
          β”‚
       JEV#2   "has this person finished their turn?"
          β”‚                        └─ false ─▢ keep holding
          β”‚ true
          β–Ό
       MicGate   mutes the mic while we are speaking
          β–Ό
      SpeakSink ─▢ TTS ─▢ speaker
          └──────────────── the mouth β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Why Turn-Taking Is the Whole Problem

A naive voice agent waits for the speech-to-text provider to announce "end of utterance", then replies. That fails in both directions, constantly:

Neither failure raises an error. Both just make the conversation feel wrong. So the system needs an actual judgement β€” and it needs two of them, because the same text gets opposite answers depending on which question you ask:

JEV#1 β€” the earJEV#2 β€” the mouth
QuestionIs there enough content to answer?Has this person finished their turn?
What it readsthe live transcript since the last submissionthe whole chunk just submitted, plus what the silence has done since
If truesubmit into the reply list, start the LLMopen the gate, start speaking
"I have three questions, the first is about indexes."true β€” that sentence is answerablefalse β€” he obviously has two more
"uh… so… the…"false β€” no contentfalse β€” still hunting for a word

Those two rows are the design. One judge cannot answer both questions, and collapsing them produces exactly the interrupting agent described above.


JEV β€” A Judgement, Not a Prompt

JEV is a judgement call made by a model through a decisions API: it takes a state string and a question with explicit criteria, and returns a calibrated scalar plus a verdict. It is not a chat completion, and the build does not treat it as one. Three consequences shape the code:

1. The question file is the policy

Both questions live in version control as authoritative files, and each contains exactly one fenced JSON block. The client reads only that block β€” everything around it is reasoning written for humans and never reaches the model. Editing the block is editing behaviour, so it is reviewed like code.

2. The examples are hand-written, never copied from the corpus

If you paste real recorded utterances into the question and then measure accuracy against those same recordings, you have measured recitation. The entire corpus stays held out as the test set, and the benchmark enforces the separation.

3. A verdict that cannot be read is an error, never a false

If the response comes back without a readable scalar, the client raises. Degrading to "not complete" instead would produce a system that never opens its mouth while reporting no error anywhere β€” the worst failure mode available.

Two further details that cost real debugging to learn:


The Reply List

The reply list is the spine of the build. It exists because speech does not arrive in neat turns, and because a crashed process must not silently lose what someone said. Every piece of state is a file in the session directory, so the conversation survives the process dying.

One utterance, three cuts, one turn

You say something, pause, add to it, pause, add again. JEV#1 passes three times, so three cuts enter the reply list. They do not become three LLM calls:

cut 1 ─┐
cut 2 ─┼──▢  one drain  ──▢  "cut one cut two cut three"  ──▢  one LLM turn
cut 3 β”€β”˜

The cuts are joined with spaces, not newlines β€” newlines nudge the model into answering each fragment separately. And when cuts arrive while the LLM is already working, they simply accumulate and are eaten in one go by the next turn. That behaviour is what the whole structure exists to produce.

The five pieces

PieceJobThe failure it prevents
Reservoirthe reply list itself β€” cuts land here, a turn takes the whole batchthe batch is taken by an atomic rename, so a cut written at the same instant cannot be lost
UserBufferthe live transcript β€” a partial overwrites the current segment, a final freezes it and starts a new oneempty partials and empty finals really do arrive from the provider; unguarded, the first wipes a half-spoken sentence and the second grows the file forever while reading perfectly normally
SilenceCounterhow long the person has actually been quietit reads the audio timeline the provider stamps on each event, not arrival time, and clamps the correction so clock drift can only under-report
SpokenLedgerwhat was really spoken, versus what was merely generatedreconciliation only ever truncates β€” over-reporting would write words into the transcript that never left the speaker, producing a lying record
VoiceDriverone pump call = exactly one turn, start to finishβ€”

A turn has exactly one ending

Earlier versions had three: finish, get interrupted, get invalidated. All of them are gone. Once a turn starts, it finishes and it is written down. If you speak while the machine is speaking, your words go into the reply list and wait β€” they do not cut it off.

This is deliberate, and it is defended at the source level, because the deleted machinery is easy to reinstate by accident: restoring one interrupt event turns no behavioural assertion red, since nobody sets it, but the next person to read the code will wire the trigger back up. So four assertions parse the source with Python's ast module and fail if the names come back.

Three places deliberately throw away what a person said

All three leave a trace, because the one thing worse than dropping a sentence is dropping it silently:

  1. Three strikes. The transcript is judged insufficient three ticks running with no growth. The "no growth" clause is the entire safeguard against cutting off someone who is still talking.
  2. The delivered watermark. JEV#1 passes on a partial, the buffer is cleared, and then that block's final arrives and writes the same sentence back into the empty buffer. Without a watermark, the same sentence gets answered twice β€” it happened in two turns out of six on a real session. The watermark cuts only a verbatim prefix, never a character of new text, and declines to cut at all when coverage drops below threshold: better to repeat a few words than to bite the head off a sentence.
  3. The hallucination gate. In one session, three seconds after the mic opened with nobody speaking, the provider emitted a confident sentence over a 0.62-second audio window. It drove a full LLM turn, a full TTS turn, and was written permanently into memory. The test is speech rate (11.31 units/sec, against a real person's fastest 4.31), not content, so the gate does not drift with the topic β€” and it defaults to letting things through when duration is missing, because losing one hallucination is cheaper than swallowing whole batches of real speech.

The Mouth

The reply does not go straight to the speaker. The first chunk is held at a gate while JEV#2 decides whether you have finished, across three bands:

Real silenceWhat happensCost
under 1.0 shold β€” do not even askfree
between the boundsask JEV#2, rate-limited to once per extra second of silenceone request
6.0 s or moreopen unconditionally β€” do not askfree

Checking the gate is free; asking JEV#2 costs money. Without the rate limit, dropping the mouth's tick interval from 1.0 s to 0.25 s would have quadrupled the bill while running perfectly and raising nothing. Once the gate opens it forwards transparently β€” only the first chunk of a turn ever waits.

MicGate then mutes the microphone for exactly as long as the machine is speaking, plus a short tail for room reverberation. Without it, the shape is: speech out of the speaker β†’ into the mic β†’ transcribed β†’ drives another turn β†’ a full LLM and a full TTS every lap, spending real money β€” and since nothing interrupts anything, it sounds like a normal conversation while it happens.

It sits inside the gate and outside the speaking sink. Outside the gate it would mute the mic while the gate is still waiting, destroying the very silence measurement the gate depends on. It releases in a finally, because one exception leaving the mic muted means permanently deaf, with nothing on screen to say so.

MicGate is not acoustic echo cancellation. It stops the machine hearing itself. It does not stop it hearing another person, or a second device playing audio. The script asks you to confirm your setup before opening the mic, every time.

β€” from the repository README

Memory

Conversation is appended to a per-peer thread file by the script, deterministically β€” the model generates text, it never writes the record. The file survives the process exiting, which also means the last run's conversation is context for the next one.

A retention window of recent blocks is injected verbatim every turn, sized by a token budget rather than a block count, with a hard floor underneath it: a single long monologue makes a pure budget compute zero blocks, and the agent then cannot remember the sentence you just said. Older blocks are compressed into a summary on a worker thread β€” running that blocking call inline jams the event loop, and for those seconds nobody consumes mic frames. The file is archived before every overwrite, because compression is lossy and that file is the only store of the verbatim words.


The Gate β€” 151 Offline Assertions

Everything above runs under a regression gate with 151 assertions: no LLM, no network, no microphone, roughly a tenth of a second, and it writes nothing outside a sandbox. It is the main reason the repository is worth reading, because nearly every assertion is a regression for a specific failure that ran, produced no error, and sounded fine.

Each assertion carries its own reasoning inline, so the file doubles as the design record β€” why barge-in was removed, why the silence reading is measured in seconds rather than ticks, why a stub's duration has to be physically plausible.

The structural technique worth stealing: where an assertion must check that a deleted mechanism has not returned, it parses the source with ast rather than grepping. The prose explaining a deletion necessarily names the deleted thing β€” so a grep would match its own documentation and stay permanently red.


Known Limits

Stated plainly, because none of these announce themselves at runtime.


Get It

GitHub: IntelligenismCommercialDevelopment-LLC/jev_voice_agent β†—

Requires Python 3.12 or newer, a microphone, and headphones. Run the offline gate first β€” it is free and takes about a tenth of a second.

Derived from the AI Operability Edition of the Intelligenism Agent Framework β€” specialized into a spoken-conversation agent built around two independent turn-taking judgements.