# Live captions need a revision history

A pilot concept: provisional text, visible corrections and separate delay measurements. Open a fictional segment history.

L00P.AI · 1.0 · 2026-09-26

https://www.l00p.ai/en/resources/series/napisy-na-zywo-potrzebuja-historii-zmian/

## 01 / Speed does not finish the job

A sentence appears on screen. Moments later, the system separates voices; later still, an editor corrects a name. To the reader, this is still the same utterance. If every update silently erases the previous version, understanding what changed becomes difficult.

We propose captions that can say “this is provisional” and then show a correction. The aim is access to the conversation, including for deaf and hard-of-hearing people. We do not assume that everyone reads at the same pace or wants the same signals.

This is a concept brief for a pilot, developed from a project note dated 5 September 2026. The example below is fictional. It receives no broadcast and does not demonstrate an operational live-caption service.

## 02 / One utterance, different kinds of knowledge

Words, speaker separation and confirmed identity are separate pieces of information. Each may arrive at a different time and each may need correction. A speaker photograph is another layer: it requires both correct attribution and permission to use it.

The proposed record preserves a stable segment identifier and a time reference in the original recording. Each revision identifies the changed field, its source and who approved it. The utterance time does not become the correction's arrival time.

Voice recognition may suggest an attribution. It should not turn that suggestion into certainty on its own. As the [guide to introducing speakers](/en/resources/series/przedstawienie-pomaga-ale-nie-potwierdza-tozsamosci/) shows, no alert does not confirm identity. A programme schedule or microphone label must also be checked against what actually happened.

## 03 / Open the successive versions

Fictional segment S-17 covers seconds 12–16 of a recording. Its first caption reads “Meeting on Thursday”. The next version adds “Speaker A”, without a name. Only listening in the scenario reveals “Tuesday”; separately, the editor confirms fictional Anna Nowak. The word change and attribution change have different reasons.

- Revision 1 / S-17 / 12–16 s: Unknown; Meeting on Thursday. Provisional transcription in the scenario.

- Revision 2 / S-17 / 12–16 s: Speaker A; Meeting on Thursday. Speaker separation; identity still unknown.

- Revision 3 / S-17 / 12–16 s: Speaker A; Meeting on Tuesday. Word corrected after fictional editorial listening.

- Revision 4 / S-17 / 12–16 s: Anna Nowak — fictional person; Meeting on Tuesday. Separate editorial attribution confirmation in the scenario.

Each panel opens independently, including by keyboard. Nothing scrolls automatically. Every version remains in the MD and JSON text. This demonstrator shows history, not a transcription algorithm; the editor and listening are parts of the invented scenario, not an audio test we performed.

## 04 / Measure delay, not just model processing

The budget starts at a defined point in the audio and ends when the reader sees a caption. Account for collecting a chunk, waiting in a queue, processing, delivery and display. Do not add parallel layers as though they always wait for one another.

These are illustrative assumptions only for four sequential stages of the first caption: 1.2 s buffering, 0.8 s queuing and processing, 0.4 s delivery and 0.1 s display. Total: 2.5 s. They are neither measurements, an agreed target nor a product promise.

![Assumed sequential-stage budget: buffering 1.2 s; queue and processing 0.8 s; delivery 0.4 s; display 0.1 s. Total 2.5 s. Example, not a measurement.](https://www.l00p.ai/wydawnictwo/napisy-na-zywo-potrzebuja-historii-zmian/opoznienie-en-v1.svg)

Original diagram, Codex / L00P.AI. Bar length shows assumed duration; scale 0–2.5 s. Later identification and corrections are excluded.

In a pilot, record the first text, later attribution and human correction times separately. State the reference point, sample count, load conditions, median and slower cases. Delay relative to the source and the gap between captions and the stream being played are different measurements. Also retain the proportion of missing segments and the count of corrections that change meaning.

## 05 / The reader must control the screen

Meaning must not depend on colour alone. “Speaker A” should remain understandable without distinguishing hues; a change needs a description. This follows [WCAG 1.4.1 — Use of Color](https://www.w3.org/WAI/WCAG22/Understanding/use-of-color.html).

The design should offer a stable view, scrolling controls and a clear return to the current position. [WCAG 2.2.2](https://www.w3.org/WAI/WCAG22/Understanding/pause-stop-hide.html) describes conditions for pausing movement and automatic updates. Pausing the view does not pause the broadcast: show when the reader is viewing an earlier passage.

Captions can also convey significant laughter, music or speaker changes when needed to understand the content. [W3C's live-caption explanation](https://www.w3.org/WAI/WCAG22/Understanding/captions-live.html) describes this. Criterion 1.2.4 concerns synchronized media; we do not present it as an automatic requirement for every audio-only radio stream or as certification of this prototype.

## 06 / What we have and have not confirmed

The 5 September note listed existing audio retrieval, transcription, speaker-separation and speaker-bank components. It also marked the revision model, caption page and visible corrections as work to build. That is the state described in a document, not present-day acceptance of an integrated service.

For this edition, we built the public revision example and budget diagram and checked the files and page. We did not test the broadcast, production latency, screen-reader operation or usability with intended readers. Vibration, photographs, a speech-only stream and source integration remain outside the demonstrator. We do not carry over historical prices, recognition thresholds or medical-benefit claims.

## 07 / A small pilot with acceptance criteria

Start with an appropriately licensed recording with an established transcript and speaker list. Only then move to a controlled transmission. Agree with participants on sharing and retention. Test a mistaken name, overlapping voices, a missing passage and connection loss.

- Can the reader recognise a provisional version and find the reason for a correction?
- Does the reading position stay stable when new text arrives?
- Is a delayed older version prevented from replacing a newer one, and a retried event from duplicating a caption?
- Does the interface work with a keyboard, magnification and the agreed assistive technologies?

Agree on criteria and acceptable delays before testing with the people who will use the captions. Their participation is needed to assess whether the solution helps. We invite a discussion about such a pilot: a description of the programme and readers' needs is enough to begin.

## 08 / Sources and AI involvement

The complete project note was read and the linked W3C explanations checked on 26 September 2026. The source document informs the concept; its historical figures were not treated as current measurements. [Example data — PL/EN JSON](/wydawnictwo/napisy-na-zywo-potrzebuja-historii-zmian/przyklad-v1.json) contains fictional revisions and arithmetic assumptions.

Codex prepared the PL/EN text, SVG and demonstrator. Review is by the author, without an independent second model. No third-party recordings, photographs or real participant data were used. Pilot results, a source change or a discovered error trigger another review. A human may withdraw the publication.
