Data model
How sources, streams, audio, turns, speakers, and contacts fit together.
Optima's data flows from recorded audio to transcript turns, and from detected voices to the people you name. This page defines each record and how they relate.
| Record | What it represents |
|---|---|
| Source | A device, iOS app, macOS app, or web app belonging to an account |
| Stream | One recording from a source |
| Audio chunk | An independently playable piece of a stream |
| Turn | A stretch of transcript text, with optional word timings |
| Speaker | A persistent detected voice |
| Contact | An account-owned person record that you create and name |
| Conversation | A revisable coherent episode with exact revision-pinned turn membership |
| Memory | A grounded claim or synthesis worth remembering |
| Suggestion | A grounded, revisable idea with its own action lifecycle and person roles |
Sources, streams, and audio chunks
A source identifies where audio comes from: a device, iOS app, macOS app, web app, or custom integration belonging to an account.
An app registers its own ios, macos, or web source for each installation. A custom integration, script, bot, or benchmark harness talking to the public API directly uses the account's single api source, which Optima creates the first time it is used; see Audio uploads. A device source is created when a device is bound to an account, and a shared_recorder source the first time an account hosts a session on a shared recorder; see Devices.
A source owns recording streams. A stream is one recording, and it is the unit you register, name, and read; a source can record several streams at the same time. A stream contains ordered, independently playable audio chunks.
Each chunk has a stable client-generated ID. Optima issues a confirmed custody receipt for the chunk before transcription work begins. Audio uploads walks through the ingest sequence. Source activity reports which sources are currently online or listening.
Turns
Processing groups audio into runs and publishes turns. A turn is an individual stretch of transcript text with optional word timings.
A turn keeps its source ID and can refer to one or more source-audio clips for playback.
Turn lists use absolute time ranges. A turn without an absolute timestamp is available by ID but does not appear in a range result.
Speakers and contacts
A detected voice is represented by a persistent speaker. The transcription provider labels speakers within each run, and that per-run label links a turn to a speaker.
Turns expose that persistent voice as speaker_id. Turns from the same recognized speaker share the ID even when no contact has been assigned; a null value means no speaker was detected.
A contact is an account-owned person record that you create and name. Processing creates speakers and turns, but it never creates contacts.
A speaker can be associated with a contact. More than one speaker can point to the same contact.
Detected and effective attribution
Each turn carries two contact values:
| Field | Meaning |
|---|---|
detected_contact_id | The contact associated with the turn's linked speaker |
contact_id | The effective contact, after any turn-specific override |
This distinction gives you two ways to correct attribution:
- Override one turn without changing other turns from the same speaker.
- Assign the speaker to a contact to change the speaker's other turns, except those with their own overrides.
Turns and contacts lists the exact patch operations.
Conversations, memories, and suggestions
A conversation organizes the timeline without owning its memories. Memories cite
exact turns and may refer to one another through one revision-pinned references
field. Conversation navigation is derived from evidence membership. A suggestion
cites pinned memory revisions and separately records action state, person roles and
date uncertainty. See Conversations and suggestions.