CEDX Speech transcribes batch and live audio, scores every job for word error rate and confidence, and runs voice agents whose interruptions and CSAT are watched as closely as their costs.
The demo's field-radio channel runs at an estimated 17.8% word error rate with 0.61 confidence — and it sits on the overview as a priced alert, not buried in a log. Hard audio is the case this product is built for.
speech.cedxsystems.com — live build
Runs on demo data. The budget card says “illustrative from the demo sample” on the screen itself.
55 hours of audio320 jobs in the sample window · batch and live
Median WER 10.9%completed jobs only — the card says so
14 voice agents23,734 sessions in 30 days · CSAT 4.1
What it is
Voice work is honest in three places.
Every transcript carries its WER
320 jobs in view with model, status, word error rate, confidence, duration and cost on each row. job_18713 — Claims intake on listen-accurate — came out at 6.8% WER and 0.83 confidence; the Support IVR job beside it ran 14.0% at 0.65. The difficult rows are not hidden: a 22.4% WER job sits flagged Warning in the same table.
WER and confidence on every completed row
Models from listen-fast to listen-accurate
3,270 minutes in the filtered view
speech — screen-2
Live streams, judged in milliseconds
48 streaming sessions with an average first byte of 527ms and an average RTF of 1.03 — wall time against audio time, printed on the header. A 13-minute field-radio stream holds 0.71 confidence; the errors, one at 0.33, are in the table with everything else.
RTF — wall time per audio time — on every session
First-byte latency averaged on the header
8 errors in the sample, not filtered out
speech — screen-3
Voice agents with manners metrics
14 voice agents, 23,734 sessions in 30 days, and the columns that make voice honest: handle time, barge-in share, tools per session and CSAT. booking.confirmation carries 6,980 sessions at 4.8 CSAT despite a 32% interrupt rate; dispatch.status-line is paused at 3.8 — which is also a decision this table supports.
Barge-in share as a first-class column
CSAT from rated sessions, 4.1 on average
A paused agent stays in the table
speech — screen-4
Product tour
Four screens, captured from the running build.
Not a mockup and not a concept deck. This is what opens at /app/speech.
speech.cedxsystems.com
01 — Overview
The bad audio is on the front page
320 jobs in the sample — 256 completed, 39 failed, 14 processing, 11 queued — with the 12.2% fail rate on a headline card. Storm week's field-radio spike is annotated on the chart, and the alerts are priced: accurate-model spend up 22% week over week at $39.40, field-radio WER risk at $22.10.
Median WER 10.9%, completed jobs only — on the card
High-risk cells: 1 of 72 scored
Budget pace 3% of $18K, labelled illustrative
02 — Transcripts
A ledger of every job
Each job row names its project and model with status, WER, confidence, duration and cost. The most expensive visible job — 18:58 of Support IVR on listen-accurate — cost $0.4552; the field-radio warnings cost $0.0000 and say why they're flagged anyway.
Chips: completed · processing · failed · queued
Duration and cost side by side on every row
CSV of the filtered view
03 — Live
Streaming, with the stopwatch showing
Forty-eight sessions with RTF, first-byte, duration, confidence and cost per row. The fastest first byte in view is 261ms; the average is 527ms. Sessions that errored keep their rows — an 0:58 stream at 0.50 confidence is right there.
14 sessions streaming right now
RTF from 0.41 to 1.71 in view
Chips: active · ended · error
04 — Voice agents
The voice is a configuration
Every agent pairs a listen model with a speak voice — listen-balanced with speak-warm, listen-stream with speak-deep — and the table watches sessions, handle time, barge-in, tools per session and CSAT. ivr.support-router averages 5:10 a call; utility.meter-report interrupts only 4% of the time.
14 agents, 13 active, 1 paused
Sessions per agent from 124 to 6,980
CSAT from 3.3 to 4.8 across the table
Who runs it
Three roles keep voice honest.
Roles, not references. We have no named customers yet, so nobody in these photographs is quoted, credited or claimed as one.
Contact-center operations
Watches the Support IVR project: 18-minute calls at 14% WER, the es-US confidence gate at 0.88, and the disconnect cluster the alert traces to packet loss.
Support IVR · 0.88 gate
Quality review
Reads transcripts against the audio: sorts by WER, opens the 22.4% rows, and decides what the glossary has to learn next.
median WER · 10.9%
Compliance and records
Keeps the retention question answerable: every job has an id, a model, a duration and a cost, exportable to CSV.
320 jobs · exportable
The shape of it
What the demo workspace actually looks like.
Every figure below is legible in the captures above. Nothing here is a projection of your estate — it is the state of the demo data.
55hof audio in the sample320 jobs · batch plus live
Voice-agent CSATrated sessions · 23,734 sessions in 30 days
4.1
Average CSAT · 4.1 of 5 across rated sessions
Barge-in share 18.0% — watched on the same header
How it runs
An audio job's life, in the order it actually happens.
01
Capture
Audio arrives as a batch job or a live stream; the demo mixes both — 55 hours across 320 jobs in the sample window.
02
Transcribe
The model choice is a stated trade: listen-fast is cheap, and listen-accurate crossed a budget alert on Aug 2 — the chart annotation says when.
03
Score
WER, confidence and gates turn quality into a queue: field radio's 17.8% estimate, the es-US IVR's 0.81 against its 0.88 gate.
04
Speak
Voice agents close the loop — a listen model and a speak voice per agent, with handle time, barge-in and CSAT watched per session.
One record
The voice of the record the estate already keeps.
Speech is not a bolt-on phone bot. Its projects in the demo — Support IVR, Claims intake, Dispatch notes — are named for the work the rest of the estate does.
Finding this out on the third call is worse for you than reading it here, and worse for us.
Speech is not generally available. What opens today is the live build running on demo data — Harbor Field Services is the software's demo workspace, not a customer.
We have no named customers to show you, so this page shows none.
The quality figures are the demo sample's, and the median WER counts completed jobs only — the card says so on the screen. Your audio will be better or worse; the demo's job is to show the mechanics.
12.2% of the sample's jobs failed. We show that rather than hide it, but it is also true that production failure handling at your scale is something a pilot has to prove.
No telephony-specific certification — nothing about call-recording compliance or regulated-industry use — is claimed on this page. What we can evidence is on the security page.
The model lineup in the captures — listen-fast through listen-accurate, the speak voices — is the demo configuration. Which models a given plan can use is not something we are claiming here.
Yes. Every screenshot is a capture of the running build and you can open the same build at /app/speech. It runs on demo data — Harbor Field Services is the demo workspace, not a customer.
What is WER, and what is a good number?
Word error rate — the share of words a transcript gets wrong, lower is better. The demo's median on completed jobs is 10.9%; its best flagged cell runs 5.1% on clean phone audio, and its worst alert is 17.8% on a field-radio channel. The honest answer is that it depends on the audio, which is why the number is printed per job.
Does it handle live audio, or only recordings?
Both. The Live screen shows 48 streaming sessions with an average first byte of 527ms and an average RTF of 1.03 — wall time against audio time — and the Transcripts screen carries the batch jobs. The overview counts both in the same 55 hours.
What happens to the jobs that fail?
They stay visible. 39 of the 320 sample jobs failed — 12.2%, on a headline card — and the failed rows keep their model, duration and cost in the table. The Live screen keeps its 8 errored sessions for the same reason.
Is Speech audited or certified?
No certification has been issued, and no call-recording or regulated-industry compliance is claimed here. What we can evidence about hosting, encryption, tenant isolation and production access is written up on the security page.
The waveforms are running. Go and listen.
Live build, demo data, no card. Then ask what your worst audio channel would score.