industryPublished July 14, 2026· Updated August 5, 2026· 8 min read

The AI Interpretation Glossary: 50+ Terms, Plainly Defined

RSI, S2S, ASR, floor language, relay, booth, ear-voice span — every term in AI interpretation and language technology, defined in plain English with examples.

Interpretation sits at the collision of two vocabularies — a century of professional interpreting, and a decade of speech AI — and vendors freely mix them. This glossary defines the terms you'll meet in product pages, procurement documents and research papers, in plain English, grouped by where they come from.

The interpreting profession

Interpretation (vs translation)

Rendering live spoken language into another language. Translation is the written counterpart. The distinction matters because live speech has no pause button: interpretation must handle accents, hesitation and interruption in real time.

Simultaneous interpretation

Interpreting while the speaker is still talking, trailing by a few seconds. The conference-booth model; cognitively the hardest mode, which is why human practitioners work in rotating pairs. See simultaneous vs consecutive.

Consecutive interpretation

Interpreting after each completed turn. Slower, more accurate, minimal equipment — the standard for legal, medical and negotiation settings, and the mode Pikka Interpret uses for meetings.

Sight translation

Orally rendering a written document on the spot — a contract read aloud in another language without preparation time.

Relay interpretation

Chaining languages through a pivot when no interpreter covers the direct pair: French → English → Japanese. AI systems mostly eliminate relays by translating every pair directly.

Floor language

The language currently being spoken to the room. In multi-language events, interpretation is organized around what the floor is speaking at any moment.

Pivot language

The intermediate language in a relay chain. Rare pairs route through a widely covered language, compounding error at each hop.

Booth

The soundproofed workspace where simultaneous interpreters work at physical events. RSI platforms are the booth's software successor.

Ear-voice span (EVS)

The lag between a simultaneous interpreter hearing a segment and speaking its rendering — typically 2–3 seconds. The human latency budget AI systems are measured against.

Chuchotage (whispered interpretation)

Simultaneous interpretation whispered to an audience of one. Pikka's personal interpreter mode is its software descendant: earphones in, the world dubbed for one listener.

Certified / sworn interpreter

An interpreter legally authorized to render for official purposes — courts, immigration, civil registries. Their rendering carries legal standing; no AI holds this in 2026.

Conference interpreter

A professional interpreter specialized in conference settings — typically simultaneous, typically paired, typically $900–$1,400 per language per day in 2026.

Community interpreter

An interpreter for public-service settings — hospitals, schools, social services — where accuracy intersects with duty of care.

RSI (remote simultaneous interpretation)

Simultaneous interpretation delivered over a cloud platform instead of physical booths — interpreters work from anywhere, listeners receive channels online. Interprefy pioneered the category.

OPI (over-the-phone interpretation)

Human interpretation by telephone, on demand — the workhorse of hospitals and call centers before video.

VRI (video remote interpretation)

Human interpretation over video — adds visual context (gesture, documents) to OPI. Common in healthcare and legal intake.

The AI stack

ASR (automatic speech recognition)

The technology that converts speech to text. Stage one of a cascade pipeline; its accuracy on accents and noise bounds everything downstream.

NMT (neural machine translation)

Text-to-text translation by neural networks — the engine behind modern text translation and stage two of a speech cascade.

TTS (text-to-speech)

Synthesizing spoken audio from text. In cascades, this is the voice you hear — a stock voice that has never met the speaker.

Speech-to-speech translation (S2S)

Translating audio directly into translated audio with one model — no text middleman. Prosody survives, and voice characteristics can ride along. Pikka Interpret runs on Google's Gemini Live Translate S2S model.

Cascade pipeline

ASR → MT → TTS chained in sequence. Modular and cheap; every hand-off adds latency and discards tone, pacing and speaker identity.

End-to-end speech translation

A single model mapping source audio to target output — the research ancestor of today's commercial S2S systems.

LLM (large language model)

The transformer-based models behind modern translation quality — context handling, terminology and disambiguation far beyond phrase-based systems.

Streaming / incremental translation

Translating audio as it arrives, chunk by chunk, instead of waiting for complete utterances. Lowers latency; risks mid-sentence revisions when the speaker changes direction.

Latency

The delay between speech and its rendering. Captions run ~1 s; simultaneous AI voice 2–4+ s (KUDO's engineers have published ~4.1 s); consecutive systems like Pikka render each completed turn immediately.

Diarization

Determining *who* spoke when — separating and labeling speakers in audio. The difference between a transcript and a meeting record.

Code-switching

Mixing languages inside one utterance — normal in multilingual communities, historically brutal for ASR and MT systems.

Glossary / terminology management

Feeding the system your product names, acronyms and domain vocabulary so they translate consistently. The single highest-leverage quality control for business interpretation.

Voice cloning

Synthesizing speech with a specific person's vocal identity. In live translation, S2S models carry the speaker's pitch, pace and timbre into the target language. See voice cloning in live translation.

Voice simulation

Microsoft's term for Teams Interpreter's version of voice cloning — transient, samples never stored, per its documentation.

Zero-shot voice cloning

Cloning a voice from a few seconds of live speech, with no enrollment recording or stored profile. How Pikka dubs every speaker from the moment they first talk.

Prosody

The melody of speech — rhythm, stress, intonation. Where meaning like sarcasm, emphasis and emotion actually lives; the first casualty of cascade pipelines.

WER (word error rate)

ASR accuracy metric: the percentage of words transcribed wrongly. Useful for benchmarking speech recognition; silent on translation quality.

BLEU

A text-translation similarity metric. Common in papers, weak for live interpretation — it measures string overlap, not whether a listener understood the meeting.

The meeting room

Floor (audio floor)

The original, untranslated speech stream of the room. Pikka listeners can toggle between the floor and their dub at any time.

Dubbing

Replacing or overlaying the original voice with translated speech. In Pikka, every speaker is dubbed into each listener's language in the speaker's cloned voice.

Language channel

One output language stream. Human events staff one interpreter (pair) per channel; AI rooms generate channels on demand per (speaker × language).

Language selector

The listener-side control for choosing a channel. Per-listener choice is what lets five listeners hear five languages from one meeting.

Captions vs subtitles

Captions transcribe the same language (accessibility); subtitles translate into another. In live AI products the terms blur — what matters is whether it's text or speech.

Translated captions

Live text translation on screen — Zoom's and Meet's native feature. Fast (~1 s), scalable, and reading-limited by design.

Live transcript

The running text record of a session. Pikka streams source and translated captions to all participants in real time.

Room code

The short join credential for a browser meeting room — Pikka uses 6-character codes, shareable as text or QR.

SFU (selective forwarding unit)

WebRTC server architecture that forwards each participant's stream without mixing — Pikka uses a mediasoup SFU so every speaker's audio arrives as its own stream.

WebRTC

The browser-native real-time audio/video standard. Why Pikka rooms need no app, plugin or download.

Echo cancellation / noise suppression

Audio pre-processing for calls. Pikka's personal mode deliberately disables both (with gain up) so the phone mic picks up the whole room, not just one voice.

Ambient capture

Recording the sound field around a device rather than a single close-up voice — the input mode behind personal interpretation.

DTX / FEC

Discontinuous transmission (sending nothing during silence) and forward error correction (rebuilding lost packets) — the Opus-codec machinery that keeps browser audio smooth on bad networks.

Buying and operations

Hour package

Buying interpretation as a block of hours (Wordly-style, commonly 12-month blocks). Fits event budgets; punishes many-small-meetings usage.

Event-day pricing

One price per language per day the room runs — Pikka's model: $249 per language per event-day, 25 seats included.

Per-seat pricing

Licensing per participant or host. Simple until guests arrive — external attendees rarely hold seats.

Pay-as-you-go (PAYG)

Usage-metered billing without commitment. KUDO's AI tier offers it alongside annual terms.

Quote-based pricing

No published price; scoped per engagement (Interprefy, KUDO enterprise). Budget discovery requires a sales cycle.

Data retention

What the vendor keeps after your session — audio, transcripts, logs — and for how long. The first privacy question in any interpretation purchase.

Zero data retention / ephemeral architecture

Processing without persistence: audio lives in RAM and is discarded at session end. Pikka's design — nothing written to disk, room close wipes state. Nothing to retain is stronger than any retention policy.

DPA (data processing agreement)

The contract governing how a vendor processes your data. Read it before any interpreted session touches regulated content.

Data residency

Where processing physically happens. Jurisdictionally decisive for EU, healthcare and government buyers.

BAA (business associate agreement)

The HIPAA instrument required before a vendor touches protected health information in the US.

ISO 27001

The international information-security certification. Interprefy and DeepL hold it; procurement teams ask for it first.

SOC 2 Type 2

An audited report on security controls operating over time. DeepL Voice's stack includes it alongside ISO 27001 and HIPAA compliance.

Interpreter marketplace

A platform for booking vetted human interpreters on demand — KUDO's is the category leader (12,000+ interpreters, 200+ languages, 12-hour lead times).

Lead time

Minimum notice to provision interpretation. Marketplaces cite hours (KUDO: 12); AI rooms cite seconds (Pikka: under a minute).

Frequently asked questions

What does RSI stand for?

Remote simultaneous interpretation — simultaneous interpretation delivered over a cloud platform instead of physical booths, letting professional interpreters work from anywhere. Interprefy pioneered the category; KUDO also operates in it.

What is speech-to-speech translation?

Translating audio directly into translated audio with a single model, with no text step in between. It preserves prosody and enables voice cloning. Pikka Interpret runs on Google's Gemini Live Translate speech-to-speech model.

What is ASR?

Automatic speech recognition — converting speech to text. It's stage one of the classic ASR → MT → TTS cascade; its accuracy on accents and noise bounds the whole pipeline.

What is the ear-voice span?

The lag between a simultaneous interpreter hearing speech and speaking its rendering — typically 2–3 seconds for professionals. It's the human latency benchmark AI systems are compared against.

What's the difference between dubbing and captions?

Captions are text on screen; dubbing is spoken audio in the target language. Dubbing carries tone and voice identity and doesn't impose reading load; captions are faster and cheaper. Pikka Interpret provides both.

What does zero data retention mean?

The vendor stores nothing from your session — audio is processed in memory and discarded. Pikka Interpret's architecture never writes audio to disk, so there is nothing to retain, leak or subpoena.

See the terms in action

Floor, dub, channels, cloned voice — one Pikka Interpret room demonstrates this whole glossary in sixty seconds.

Host a Meeting

Written by the Pikka Interpret team.We build real-time AI interpretation for meetings — every speaker dubbed into each listener's language, in the speaker's own cloned voice, in the browser. Facts about third-party products come from their public materials as of the updated date above; linked sources are provided for claims from studies and vendor documentation.