The AI Interpretation Glossary: 50+ Terms, Plainly Defined
RSI, S2S, ASR, floor language, relay, booth, ear-voice span — every term in AI interpretation and language technology, defined in plain English with examples.
Interpretation sits at the collision of two vocabularies — a century of professional interpreting, and a decade of speech AI — and vendors freely mix them. This glossary defines the terms you'll meet in product pages, procurement documents and research papers, in plain English, grouped by where they come from.
The interpreting profession
Interpretation (vs translation)
Rendering live spoken language into another language. Translation is the written counterpart. The distinction matters because live speech has no pause button: interpretation must handle accents, hesitation and interruption in real time.
Simultaneous interpretation
Interpreting while the speaker is still talking, trailing by a few seconds. The conference-booth model; cognitively the hardest mode, which is why human practitioners work in rotating pairs. See simultaneous vs consecutive.
Consecutive interpretation
Interpreting after each completed turn. Slower, more accurate, minimal equipment — the standard for legal, medical and negotiation settings, and the mode Pikka Interpret uses for meetings.
Sight translation
Orally rendering a written document on the spot — a contract read aloud in another language without preparation time.
Relay interpretation
Chaining languages through a pivot when no interpreter covers the direct pair: French → English → Japanese. AI systems mostly eliminate relays by translating every pair directly.
Floor language
The language currently being spoken to the room. In multi-language events, interpretation is organized around what the floor is speaking at any moment.
Pivot language
The intermediate language in a relay chain. Rare pairs route through a widely covered language, compounding error at each hop.
Booth
The soundproofed workspace where simultaneous interpreters work at physical events. RSI platforms are the booth's software successor.
Ear-voice span (EVS)
The lag between a simultaneous interpreter hearing a segment and speaking its rendering — typically 2–3 seconds. The human latency budget AI systems are measured against.
Chuchotage (whispered interpretation)
Simultaneous interpretation whispered to an audience of one. Pikka's personal interpreter mode is its software descendant: earphones in, the world dubbed for one listener.
Certified / sworn interpreter
An interpreter legally authorized to render for official purposes — courts, immigration, civil registries. Their rendering carries legal standing; no AI holds this in 2026.
Conference interpreter
A professional interpreter specialized in conference settings — typically simultaneous, typically paired, typically $900–$1,400 per language per day in 2026.
Community interpreter
An interpreter for public-service settings — hospitals, schools, social services — where accuracy intersects with duty of care.
RSI (remote simultaneous interpretation)
Simultaneous interpretation delivered over a cloud platform instead of physical booths — interpreters work from anywhere, listeners receive channels online. Interprefy pioneered the category.
OPI (over-the-phone interpretation)
Human interpretation by telephone, on demand — the workhorse of hospitals and call centers before video.
VRI (video remote interpretation)
Human interpretation over video — adds visual context (gesture, documents) to OPI. Common in healthcare and legal intake.
The AI stack
ASR (automatic speech recognition)
The technology that converts speech to text. Stage one of a cascade pipeline; its accuracy on accents and noise bounds everything downstream.
NMT (neural machine translation)
Text-to-text translation by neural networks — the engine behind modern text translation and stage two of a speech cascade.
TTS (text-to-speech)
Synthesizing spoken audio from text. In cascades, this is the voice you hear — a stock voice that has never met the speaker.
Speech-to-speech translation (S2S)
Translating audio directly into translated audio with one model — no text middleman. Prosody survives, and voice characteristics can ride along. Pikka Interpret runs on Google's Gemini Live Translate S2S model.
Cascade pipeline
ASR → MT → TTS chained in sequence. Modular and cheap; every hand-off adds latency and discards tone, pacing and speaker identity.
End-to-end speech translation
A single model mapping source audio to target output — the research ancestor of today's commercial S2S systems.
LLM (large language model)
The transformer-based models behind modern translation quality — context handling, terminology and disambiguation far beyond phrase-based systems.
Streaming / incremental translation
Translating audio as it arrives, chunk by chunk, instead of waiting for complete utterances. Lowers latency; risks mid-sentence revisions when the speaker changes direction.
Latency
The delay between speech and its rendering. Captions run ~1 s; simultaneous AI voice 2–4+ s (KUDO's engineers have published ~4.1 s); consecutive systems like Pikka render each completed turn immediately.
Diarization
Determining *who* spoke when — separating and labeling speakers in audio. The difference between a transcript and a meeting record.
Code-switching
Mixing languages inside one utterance — normal in multilingual communities, historically brutal for ASR and MT systems.
Glossary / terminology management
Feeding the system your product names, acronyms and domain vocabulary so they translate consistently. The single highest-leverage quality control for business interpretation.
Voice cloning
Synthesizing speech with a specific person's vocal identity. In live translation, S2S models carry the speaker's pitch, pace and timbre into the target language. See voice cloning in live translation.
Voice simulation
Microsoft's term for Teams Interpreter's version of voice cloning — transient, samples never stored, per its documentation.
Zero-shot voice cloning
Cloning a voice from a few seconds of live speech, with no enrollment recording or stored profile. How Pikka dubs every speaker from the moment they first talk.
Prosody
The melody of speech — rhythm, stress, intonation. Where meaning like sarcasm, emphasis and emotion actually lives; the first casualty of cascade pipelines.
WER (word error rate)
ASR accuracy metric: the percentage of words transcribed wrongly. Useful for benchmarking speech recognition; silent on translation quality.
BLEU
A text-translation similarity metric. Common in papers, weak for live interpretation — it measures string overlap, not whether a listener understood the meeting.
The meeting room
Floor (audio floor)
The original, untranslated speech stream of the room. Pikka listeners can toggle between the floor and their dub at any time.
Dubbing
Replacing or overlaying the original voice with translated speech. In Pikka, every speaker is dubbed into each listener's language in the speaker's cloned voice.
Language channel
One output language stream. Human events staff one interpreter (pair) per channel; AI rooms generate channels on demand per (speaker × language).
Language selector
The listener-side control for choosing a channel. Per-listener choice is what lets five listeners hear five languages from one meeting.
Captions vs subtitles
Captions transcribe the same language (accessibility); subtitles translate into another. In live AI products the terms blur — what matters is whether it's text or speech.
Translated captions
Live text translation on screen — Zoom's and Meet's native feature. Fast (~1 s), scalable, and reading-limited by design.
Live transcript
The running text record of a session. Pikka streams source and translated captions to all participants in real time.
Room code
The short join credential for a browser meeting room — Pikka uses 6-character codes, shareable as text or QR.
SFU (selective forwarding unit)
WebRTC server architecture that forwards each participant's stream without mixing — Pikka uses a mediasoup SFU so every speaker's audio arrives as its own stream.
WebRTC
The browser-native real-time audio/video standard. Why Pikka rooms need no app, plugin or download.
Echo cancellation / noise suppression
Audio pre-processing for calls. Pikka's personal mode deliberately disables both (with gain up) so the phone mic picks up the whole room, not just one voice.
Ambient capture
Recording the sound field around a device rather than a single close-up voice — the input mode behind personal interpretation.
DTX / FEC
Discontinuous transmission (sending nothing during silence) and forward error correction (rebuilding lost packets) — the Opus-codec machinery that keeps browser audio smooth on bad networks.
Buying and operations
Hour package
Buying interpretation as a block of hours (Wordly-style, commonly 12-month blocks). Fits event budgets; punishes many-small-meetings usage.
Event-day pricing
One price per language per day the room runs — Pikka's model: $249 per language per event-day, 25 seats included.
Per-seat pricing
Licensing per participant or host. Simple until guests arrive — external attendees rarely hold seats.
Pay-as-you-go (PAYG)
Usage-metered billing without commitment. KUDO's AI tier offers it alongside annual terms.
Quote-based pricing
No published price; scoped per engagement (Interprefy, KUDO enterprise). Budget discovery requires a sales cycle.
Data retention
What the vendor keeps after your session — audio, transcripts, logs — and for how long. The first privacy question in any interpretation purchase.
Zero data retention / ephemeral architecture
Processing without persistence: audio lives in RAM and is discarded at session end. Pikka's design — nothing written to disk, room close wipes state. Nothing to retain is stronger than any retention policy.
DPA (data processing agreement)
The contract governing how a vendor processes your data. Read it before any interpreted session touches regulated content.
Data residency
Where processing physically happens. Jurisdictionally decisive for EU, healthcare and government buyers.
BAA (business associate agreement)
The HIPAA instrument required before a vendor touches protected health information in the US.
ISO 27001
The international information-security certification. Interprefy and DeepL hold it; procurement teams ask for it first.
SOC 2 Type 2
An audited report on security controls operating over time. DeepL Voice's stack includes it alongside ISO 27001 and HIPAA compliance.
Interpreter marketplace
A platform for booking vetted human interpreters on demand — KUDO's is the category leader (12,000+ interpreters, 200+ languages, 12-hour lead times).
Lead time
Minimum notice to provision interpretation. Marketplaces cite hours (KUDO: 12); AI rooms cite seconds (Pikka: under a minute).
Frequently asked questions
What does RSI stand for?
Remote simultaneous interpretation — simultaneous interpretation delivered over a cloud platform instead of physical booths, letting professional interpreters work from anywhere. Interprefy pioneered the category; KUDO also operates in it.
What is speech-to-speech translation?
Translating audio directly into translated audio with a single model, with no text step in between. It preserves prosody and enables voice cloning. Pikka Interpret runs on Google's Gemini Live Translate speech-to-speech model.
What is ASR?
Automatic speech recognition — converting speech to text. It's stage one of the classic ASR → MT → TTS cascade; its accuracy on accents and noise bounds the whole pipeline.
What is the ear-voice span?
The lag between a simultaneous interpreter hearing speech and speaking its rendering — typically 2–3 seconds for professionals. It's the human latency benchmark AI systems are compared against.
What's the difference between dubbing and captions?
Captions are text on screen; dubbing is spoken audio in the target language. Dubbing carries tone and voice identity and doesn't impose reading load; captions are faster and cheaper. Pikka Interpret provides both.
What does zero data retention mean?
The vendor stores nothing from your session — audio is processed in memory and discarded. Pikka Interpret's architecture never writes audio to disk, so there is nothing to retain, leak or subpoena.
See the terms in action
Floor, dub, channels, cloned voice — one Pikka Interpret room demonstrates this whole glossary in sixty seconds.
Host a MeetingWritten by the Pikka Interpret team.We build real-time AI interpretation for meetings — every speaker dubbed into each listener's language, in the speaker's own cloned voice, in the browser. Facts about third-party products come from their public materials as of the updated date above; linked sources are provided for claims from studies and vendor documentation.