EUREKADEV / CODING / COMPUTER SCIENCE + AI

Hear the moment.
Inspect the evidence.

SpeechTrace compares a short speech recording with a reference, aligns the acoustics and transcript, and points to specific delivery differences. It measures reference similarity rather than a speaker's talent or the quality of their argument.

Four-minute judge walkthrough

  1. Choose an unchanged control. It should show no flagged regions. Bundled analysis is labeled Precomputed example.
  2. Choose a severe local-volume edit. Inspect the flagged interval, the measured difference and the words overlapping it. Play the paired recordings at that moment.
  3. Compare the global-gain control: a whole-recording volume difference should not become a local delivery flaw.
  4. Choose an inserted pause, then a rushed example. Inspect the pacing explanation and the energy / pacing / contour breakdown.
  5. Press Reanalyze recordings. This invokes the actual Python API and pretrained Vosk model. Upload your own two short clips and their shared transcript to use the same pipeline.
  6. Export JSON to inspect scores, boundaries, intervals and explanations. Changing input discards stale results.

October 7 rehearsal extension

Write a rehearsal intention and mark each region as intentional, worth practicing, or for later review. Download plain-text notes with your decisions. The score stays unchanged; changing the input or reanalyzing clears old notes.

Technical contribution

FFT and mel cepstral features feed a monotone dynamic-time-warping alignment. Vosk/Kaldi provides word boundaries using the supplied transcript. Ordered transcript matching addresses repeated words. Relative energy, pitch contour and local pacing yield an explicit, reproducible rubric.

Evidence and limits

The 14 pairs use two public-domain excerpts from one historical speaker. Four unchanged / global-gain controls have no flags. Reported mean interval IoU is 0.975 for moderate/severe volume edits and 0.815 for inserted pauses. Both rushed cases produce pacing flags. These are deliberately modified recordings with known edits, not a human speech-quality benchmark.

One speaker cannot establish cross-speaker accuracy or fairness. Estimated word timestamps need manual validation. Differences may be artistically intentional. Listen before interpreting a flag.

Next study, clearly proposed

Recruit consenting speakers; record matched texts; have blinded raters annotate delivery differences; separate speakers between calibration and evaluation; report interval overlap, false positives and uncertainty by recording conditions. This work has not been performed.

Origin, privacy and AI disclosure

Built October 6, 2026 for Multimodal AI Track C and previously submitted there. EurekaDev permits cross-posting; that origin is disclosed. The October 7 work adds this category-specific guide and a rehearsal annotation/export workflow. Codex assisted with code, data transformations, debugging and documentation. Vosk is pretrained; this project did not train its acoustic model. Uploaded audio and transcripts are processed for the request and are not persisted by the application. Hosting request metadata still exists.