A trainer in a foam-lined sound booth wearing a headset, a laptop showing two audio waveforms beside a hand-written “No phone” sign.

CEDX Speech · AI

Voice, measured
word by word.

CEDX Speech transcribes batch and live audio, scores every job for word error rate and confidence, and runs voice agents whose interruptions and CSAT are watched as closely as their costs.

The demo's field-radio channel runs at an estimated 17.8% word error rate with 0.61 confidence — and it sits on the overview as a priced alert, not buried in a log. Hard audio is the case this product is built for.

speech.cedxsystems.com — live build
CEDX Speech overview: job, minute, WER, latency and fail-rate cards, a minutes-versus-jobs chart, jobs by status, budget pace and root-caused alerts.

Runs on demo data. The budget card says “illustrative from the demo sample” on the screen itself.

55 hours of audio320 jobs in the sample window · batch and live
Median WER 10.9%completed jobs only — the card says so
14 voice agents23,734 sessions in 30 days · CSAT 4.1

What it is

Voice work is honest in three places.

Every transcript carries its WER

320 jobs in view with model, status, word error rate, confidence, duration and cost on each row. job_18713 — Claims intake on listen-accurate — came out at 6.8% WER and 0.83 confidence; the Support IVR job beside it ran 14.0% at 0.65. The difficult rows are not hidden: a 22.4% WER job sits flagged Warning in the same table.

  • WER and confidence on every completed row
  • Models from listen-fast to listen-accurate
  • 3,270 minutes in the filtered view
speech — screen-2
Every transcript carries its WER

Live streams, judged in milliseconds

48 streaming sessions with an average first byte of 527ms and an average RTF of 1.03 — wall time against audio time, printed on the header. A 13-minute field-radio stream holds 0.71 confidence; the errors, one at 0.33, are in the table with everything else.

  • RTF — wall time per audio time — on every session
  • First-byte latency averaged on the header
  • 8 errors in the sample, not filtered out
speech — screen-3
Live streams, judged in milliseconds

Voice agents with manners metrics

14 voice agents, 23,734 sessions in 30 days, and the columns that make voice honest: handle time, barge-in share, tools per session and CSAT. booking.confirmation carries 6,980 sessions at 4.8 CSAT despite a 32% interrupt rate; dispatch.status-line is paused at 3.8 — which is also a decision this table supports.

  • Barge-in share as a first-class column
  • CSAT from rated sessions, 4.1 on average
  • A paused agent stays in the table
speech — screen-4
Voice agents with manners metrics

Product tour

Four screens, captured from the running build.

Not a mockup and not a concept deck. This is what opens at /app/speech.

speech.cedxsystems.com
CEDX Speech Overview screen.CEDX Speech Transcripts screen.CEDX Speech Live screen.CEDX Speech Voice agents screen.

01 — Overview

The bad audio is on the front page

320 jobs in the sample — 256 completed, 39 failed, 14 processing, 11 queued — with the 12.2% fail rate on a headline card. Storm week's field-radio spike is annotated on the chart, and the alerts are priced: accurate-model spend up 22% week over week at $39.40, field-radio WER risk at $22.10.

  • Median WER 10.9%, completed jobs only — on the card
  • High-risk cells: 1 of 72 scored
  • Budget pace 3% of $18K, labelled illustrative

02 — Transcripts

A ledger of every job

Each job row names its project and model with status, WER, confidence, duration and cost. The most expensive visible job — 18:58 of Support IVR on listen-accurate — cost $0.4552; the field-radio warnings cost $0.0000 and say why they're flagged anyway.

  • Chips: completed · processing · failed · queued
  • Duration and cost side by side on every row
  • CSV of the filtered view

03 — Live

Streaming, with the stopwatch showing

Forty-eight sessions with RTF, first-byte, duration, confidence and cost per row. The fastest first byte in view is 261ms; the average is 527ms. Sessions that errored keep their rows — an 0:58 stream at 0.50 confidence is right there.

  • 14 sessions streaming right now
  • RTF from 0.41 to 1.71 in view
  • Chips: active · ended · error

04 — Voice agents

The voice is a configuration

Every agent pairs a listen model with a speak voice — listen-balanced with speak-warm, listen-stream with speak-deep — and the table watches sessions, handle time, barge-in, tools per session and CSAT. ivr.support-router averages 5:10 a call; utility.meter-report interrupts only 4% of the time.

  • 14 agents, 13 active, 1 paused
  • Sessions per agent from 124 to 6,980
  • CSAT from 3.3 to 4.8 across the table

Who runs it

Three roles keep voice honest.

Roles, not references. We have no named customers yet, so nobody in these photographs is quoted, credited or claimed as one.

Contact-center operations

Watches the Support IVR project: 18-minute calls at 14% WER, the es-US confidence gate at 0.88, and the disconnect cluster the alert traces to packet loss.

Support IVR · 0.88 gate

Quality review

Reads transcripts against the audio: sorts by WER, opens the 22.4% rows, and decides what the glossary has to learn next.

median WER · 10.9%

Compliance and records

Keeps the retention question answerable: every job has an id, a model, a duration and a cost, exportable to CSV.

320 jobs · exportable

The shape of it

What the demo workspace actually looks like.

Every figure below is legible in the captures above. Nothing here is a projection of your estate — it is the state of the demo data.

55hof audio in the sample320 jobs · batch plus live
10.9%median WER, completed jobsbest alert 5.1% · worst 17.8%
527msaverage first bytelive streams · average RTF 1.03
18.0%average barge-in share23,734 voice-agent sessions in 30 days
Root-caused alerts, priced on screensymptom → cause → fix · last 30 days
  • Accurate-model spend +22% week over week — crossed the budget alert on Aug 2$39.40
  • Claims accurate cell, low risk — listen-accurate on phone: WER 5.1%, confidence 0.93$39.40
  • Field radio WER risk elevated — listen-fast on the radio channel, est. WER 17.8% · conf 0.61$22.10
  • es-US IVR confidence under gate — 0.81 average against the 0.88 gate$4.20
Jobs by status256 completed of 320 in the sample
  • Completed · 256 jobs
  • The rest · 64 — 39 failed, 14 processing, 11 queued
Voice-agent CSATrated sessions · 23,734 sessions in 30 days
4.1
  • Average CSAT · 4.1 of 5 across rated sessions
  • Barge-in share 18.0% — watched on the same header

How it runs

An audio job's life, in the order it actually happens.

01

Capture

Audio arrives as a batch job or a live stream; the demo mixes both — 55 hours across 320 jobs in the sample window.

02

Transcribe

The model choice is a stated trade: listen-fast is cheap, and listen-accurate crossed a budget alert on Aug 2 — the chart annotation says when.

03

Score

WER, confidence and gates turn quality into a queue: field radio's 17.8% estimate, the es-US IVR's 0.81 against its 0.88 gate.

04

Speak

Voice agents close the loop — a listen model and a speak voice per agent, with handle time, barge-in and CSAT watched per session.

One record

The voice of the record
the estate already keeps.

Speech is not a bolt-on phone bot. Its projects in the demo — Support IVR, Claims intake, Dispatch notes — are named for the work the rest of the estate does.

All 132 applications

Limits

What Speech does not do yet.

Finding this out on the third call is worse for you than reading it here, and worse for us.

Start

Open it before you talk to anyone.

Pilot

Your audio, your gates

  • Everything in Try
  • Channel and audio-profile review
  • WER and confidence-gate setup
  • Glossary workshop
Talk to sales

Estate

Speech with the rest of it

  • Speech with Desk, AI Gateway and Copilot
  • One identity, one bill
  • CEDX delivery
Book an estate map

Questions

Before you pilot Speech.

Is the software on this page real?

Yes. Every screenshot is a capture of the running build and you can open the same build at /app/speech. It runs on demo data — Harbor Field Services is the demo workspace, not a customer.

What is WER, and what is a good number?

Word error rate — the share of words a transcript gets wrong, lower is better. The demo's median on completed jobs is 10.9%; its best flagged cell runs 5.1% on clean phone audio, and its worst alert is 17.8% on a field-radio channel. The honest answer is that it depends on the audio, which is why the number is printed per job.

Does it handle live audio, or only recordings?

Both. The Live screen shows 48 streaming sessions with an average first byte of 527ms and an average RTF of 1.03 — wall time against audio time — and the Transcripts screen carries the batch jobs. The overview counts both in the same 55 hours.

What happens to the jobs that fail?

They stay visible. 39 of the 320 sample jobs failed — 12.2%, on a headline card — and the failed rows keep their model, duration and cost in the table. The Live screen keeps its 8 errored sessions for the same reason.

Is Speech audited or certified?

No certification has been issued, and no call-recording or regulated-industry compliance is claimed here. What we can evidence about hosting, encryption, tenant isolation and production access is written up on the security page.

The waveforms are running. Go and listen.

Live build, demo data, no card. Then ask what your worst audio channel would score.