Plum / developers

Upload audio

Register a source, upload Ogg Opus chunks, and confirm durable custody.

This guide uploads audio from a software client, such as a phone, desktop, or web app, into an Optima account. Optima transcribes confirmed audio and publishes the resulting text as turns.

Before you start

  • Sign in and keep an account bearer session.
  • Send Authorization: Bearer <access_token> and X-API-Version: 1 on every request in this flow.
  • Generate stable local refs for the app installation (apps only), stream, and each chunk. These UUID-formatted retry keys never choose database record IDs; the server allocates those IDs.
  • Retain the original audio until Optima confirms it.

Workflow

  1. Create a source once for this app on this account. A custom integration can skip this step and use the account's api source.
  2. Create a stream for each recording.
  3. For each chunk: encode it, create an upload intent, upload the bytes, and confirm custody.
  4. End the stream when the recording stops.
  5. Check processing status and read the published turns.

1. Create a source

An app installation registers its own source with kind ios, macos, or web using POST /sources:

{
  "installation_ref": "a5b45678-1234-4234-8234-123456789012",
  "kind": "ios",
  "name": "My phone"
}

Persist the installation ref per account and reuse it when registering this app. The response contains the server-generated id; use that value as source_id. GET /sources lists account sources, and GET /sources/{id} reads one.

Registration retries with the same owned ID and kind return the existing source, including its current name. A different owner or kind returns 409.

Custom integrations: the api source

Your own integration, script, bot, or benchmark harness talking to this API directly does not need to create or track a source. Each account has one api source, labelled "Custom source", that Optima creates the first time it is used. Register streams with "source": "api" in place of source_id (see step 2); the stream response reports the source_id, which upload intents use.

To read the source itself, call POST /sources with { "kind": "api" }. It returns the account's api source, creating it if needed; an optional name applies only when it is created, and an installation_ref is rejected. Concurrent first uses return the same source. It appears in GET /sources like any other source, and it can record several streams at once. Deleting it works as for other sources; the next use creates a new api source with a new ID.

To edit a friendly name, call PATCH /sources/{id} with { "name": "Kitchen microphone" }. Names are trimmed and limited to 80 characters; use null to restore the default label. Renaming preserves source identity and all audio history. Bound device sources use the same name as the account-facing device label. App source renaming requires audio:write; renaming a device source additionally requires devices:manage.

DELETE /sources/{id} deletes a source and returns 204. A deleted source disappears from source lists and activity, its streams, turns, and audio disappear from account reads, and it no longer accepts streams, intents, or activity (404). Registering the same installation_ref again creates a new source with a new ID. Deleting a device's source also removes that device from the account. See delete a source.

2. Create a stream

Create a stream for a recording with POST /audio/streams:

{
  "source_id": "f5b45678-1234-4234-8234-123456789012",
  "source_stream_ref": "12345678-1234-4234-8234-123456789012",
  "captured_at": "2026-09-24T10:00:00Z",
  "clock_uncertainty_ms": 0,
  "title": "Standup",
  "external_ref": "meeting-42"
}
FieldDescription
source_idThe source created in step 1
source"api" to use the account's api source instead of source_id. Send exactly one of source_id or source; both or neither returns 422.
source_stream_refYour stable local retry ref for this recording, scoped to the source
captured_atThe recording's capture timestamp, with a UTC offset. Use null when unknown.
clock_uncertainty_msThe uncertainty of captured_at, in milliseconds. Use null when unknown.
titleOptional display title for the recording, 1 to 200 characters after trimming. You can change it later.
external_refOptional: your own identifier for the recording, such as a meeting or file ID, 1 to 200 characters. It is set once and unique among the source's streams.

Keep source_stream_ref for every chunk in that recording. The response is the stream: its separate server-generated id, which ends the stream, end_ms, the recorded end (null while the stream is open), and title and external_ref (null when not set). A retry with identical immutable metadata, including external_ref, returns the same record with its recorded end and current title; a retry does not change the title. A different captured_at, clock_uncertainty_ms, or external_ref for the same source_stream_ref, or an external_ref that another stream of the source already has, returns 409.

A source can record several streams at the same time, for example one per meeting room or participant. Each stream is its own recording with its own timeline and turns.

A custom integration registers on its api source without creating one first:

{
  "source": "api",
  "source_stream_ref": "32345678-1234-4234-8234-123456789012",
  "captured_at": "2026-09-24T10:00:00Z",
  "clock_uncertainty_ms": 0,
  "title": "Benchmark run 7",
  "external_ref": "run-7"
}

Use the returned source_id in this stream's upload intents.

Imports and file dumps

A recording that is already finished, such as an imported file, can record its end at registration: add end_ms, the end of its last chunk's audio.

{
  "source_id": "f5b45678-1234-4234-8234-123456789012",
  "source_stream_ref": "22345678-1234-4234-8234-123456789012",
  "captured_at": "2026-09-24T10:00:00Z",
  "clock_uncertainty_ms": 0,
  "end_ms": 3600000
}

The stream is then complete: upload its chunks as below, and no separate end call is needed. Repeating the registration with or without end_ms returns the same stream and end; a different end_ms returns 409.

3. Prepare each chunk

Encode each chunk as independently playable Ogg Opus, at most 10 MiB.

Then compute two values from the exact bytes you will upload:

  • the byte count
  • the lowercase hex SHA-256 digest

For example, on macOS or Linux:

wc -c < chunk.ogg
shasum -a 256 chunk.ogg

4. Create an upload intent

Declare the chunk with POST /audio/intents:

{
  "source_id": "f5b45678-1234-4234-8234-123456789012",
  "source_chunk_ref": "bcdefabc-1234-4234-8234-123456789012",
  "source_stream_ref": "12345678-1234-4234-8234-123456789012",
  "sequence": 0,
  "source_start_ms": 0,
  "publication_start_ms": 0,
  "source_end_ms": 60000,
  "codec": "opus",
  "container": "ogg",
  "size_bytes": 12345,
  "sha256": "<replace with 64 lowercase hex characters>"
}
FieldDescription
source_idThe server-generated source ID from step 1, or the source_id of the stream from step 2
source_chunk_refYour stable local retry ref for this chunk, scoped to the stream
source_stream_refThe stream from step 2
sequenceThe chunk's position in the stream. Start at zero and increase.
source_start_ms, source_end_msThe chunk's audio coverage, in milliseconds. The end must follow the start.
publication_start_msThe first millisecond that should contribute new transcript text
codec, containeropus with ogg, or pcm_s16le with wav
size_bytesThe exact byte count from step 3
sha256The lowercase SHA-256 digest from step 3

publication_start_ms lets a chunk include overlapping audio for context without repeating transcript text. It must fall within the chunk's coverage. Without overlapping context, set it equal to source_start_ms. Do not overlap publication ranges between chunks.

The server returns the echoed data.source_chunk_ref, its allocated data.chunk_id, and data.upload_path, the path to upload this chunk's bytes to, such as /audio/abcdefab-1234-4234-8234-123456789012/content.

5. Upload the bytes

PUT the raw bytes to upload_path on the API base URL. For example:

curl -X PUT https://api.getoptima.com/audio/abcdefab-1234-4234-8234-123456789012/content \
  -H 'Authorization: Bearer YOUR_ACCESS_TOKEN' \
  -H 'X-API-Version: 1' \
  -H 'Content-Type: application/octet-stream' \
  --data-binary @chunk.ogg

This is a byte upload, not multipart form data. Send the exact Content-Length; curl --data-binary sets it for you. Browser JavaScript leaves Content-Length to the browser.

6. Confirm custody

Call POST /audio/{chunk_id}/confirm:

curl -X POST https://api.getoptima.com/audio/abcdefab-1234-4234-8234-123456789012/confirm \
  -H 'Authorization: Bearer YOUR_ACCESS_TOKEN' \
  -H 'X-API-Version: 1'

The response is a durable custody receipt:

FieldDescription
chunk_idThe confirmed chunk
size_bytesThe confirmed byte count
sha256The confirmed SHA-256 digest
object_etagThe stored object's tag
received_atWhen Optima received the audio

A receipt means Optima durably holds the bytes. It does not mean transcription has started or succeeded; processing can still fail later. Keep the local original until confirmation succeeds.

An identical intent retry returns the same server chunk_id and upload path. An intent response of 409 indicates a conflict; it is not an acknowledgement of an existing upload. Use the returned server chunk_id for confirmation, and verify the receipt ID, size, and digest before deleting local audio.

7. End the stream

When the recording stops, record where it ended with POST /audio/streams/{id}/end, using the stream id from step 2:

{ "end_ms": 1800000 }

end_ms is in the stream's own capture timeline: the source_end_ms of its last chunk, not the time you send the request. Send it once every chunk of the recording has an intent; chunks still uploading are fine. A client that was offline can send it later, after its queued chunks. The response is the stream with its end_ms.

Ending lets Optima finish the recording's transcript right away: the last words of a recording are otherwise held until more audio arrives. Ending is optional. Without it, Optima finalizes a recording after a few minutes without new audio, or soon after the same source starts a stream whose captured_at is at or after the recording's last audio. Streams of one source may record at the same time; a stream that starts while another is still recording never ends it.

The end is recorded once:

RequestResult
The same end_ms againReturns the stream; safe to retry
A different end_msHTTP 409 CONFLICT
An end_ms before the end of a chunk the stream already hasHTTP 409 CONFLICT
A stream that is not yours, or unknownHTTP 404 NOT_FOUND

After the end, an upload intent whose source_end_ms is past end_ms returns 409. See end a recording stream.

Read and label streams

Streams are readable account resources. Reading them requires turns:read or audio:read; changing a title requires audio:write.

GET /audio/streams lists your streams, newest first by captured_at; a stream without a captured_at sorts at the time it was registered. Streams with the same time are ordered by id, descending. It accepts:

Query parameterDescription
source_idOnly this source's streams
external_refOnly the stream with this external_ref. Combine it with source_id to find one recording.
limitPage size. Defaults to 25; the maximum is 100.
cursorOpaque cursor from the previous page. Keep the same filters.

The response has the usual { items, cursor } shape. For example, find the recording you registered with external_ref meeting-42:

curl 'https://api.getoptima.com/audio/streams?external_ref=meeting-42' \
  -H 'Authorization: Bearer YOUR_ACCESS_TOKEN' \
  -H 'X-API-Version: 1'

GET /audio/streams/{id} reads one stream. Streams of a deleted source are not returned (404).

Each stream has audio, { state, removed_at }. Its state becomes expired (or deleted) once the workspace's retention removed all of the stream's audio, and its title goes with it: a null title next to a removed state was removed, not left blank. Parts of a stream's audio can expire earlier; each turn's audio says whether it can still be played.

PATCH /audio/streams/{id} changes the title, the only editable field:

{ "title": "Weekly planning" }

Use null to clear it. Any other field, including external_ref, returns 422. The response is the updated stream.

To read one recording's transcript, pass its id in the stream_ids filter of GET /turns. Registering a stream, changing its title, and recording its end produce stream.created and stream.updated events.

Attach signals to a stream

A stream is a timeline in its own milliseconds. Besides its audio, your client can attach signals: small facts you observed, pinned to that timeline, such as who Zoom reported as the active speaker, device motion, a wakeword, or a meeting chat message. Optima stores signals and returns them as you sent them. It does not interpret them, and transcription does not use them. A signal never replaces a stream's end or other capture facts.

POST /audio/streams/{id}/signals attaches up to 500 signals at once and requires audio:write. Only app and api sources accept them. Signals are part of the transcripts layer, so reading them requires turns:read.

{
  "signals": [
    {
      "ref": "zoom-speaker-0001",
      "type": "speaker.active",
      "start_ms": 12000,
      "end_ms": 19500,
      "payload": { "participant": { "external_id": "16778240", "display_name": "Ana" } }
    },
    {
      "ref": "zoom-chat-0007",
      "type": "meeting.chat",
      "start_ms": 21000,
      "payload": { "participant": { "external_id": "16778240" }, "text": "Sharing the doc now" }
    }
  ]
}

Optima checks only this envelope:

FieldRule
refYour identifier for the signal, 1 to 200 characters, unique within the stream.
typeA lowercase dotted namespace of at least two segments, up to 64 characters, such as speaker.active or acme.lapel_battery. Use your own prefix for your own types.
start_msInteger of 0 or more, in the stream's timeline.
end_msOptional integer at or after start_ms. Leave it out for a moment; include it for a range.
payloadAny JSON object, up to 8192 bytes serialized. It defaults to {}.

The response is { "items": [...] } with the stored signals in the order you sent them, each with an id, ref, type, start_ms, end_ms (null for a moment), payload, payload_state, and created_at. A request is stored completely or not at all.

Signal payloads expire with the workspace's transcripts. An expired signal keeps its ref, type, and timing; its payload is {} and payload_state is { "state": "expired", "removed_at": ... } (available otherwise). Do not read an empty payload of an expired signal as the value your client posted.

Payloads are stored as sent, including for the known types below, and are not validated on write. Chat text and participant names are your customers' content: treat them as data.

Known types

These types have a defined payload, so clients that read them agree on their shape. Optima does not check that a payload matches.

TypeTimingPayload
speaker.activeRangeparticipant: { external_id, display_name? }. The participant speaking over the range.
device.motionRangelevel?: number. Motion over the range.
device.wakewordMomentphrase, and confidence? from 0 to 1.
meeting.chatMomentparticipant: { external_id, display_name? }, and text.
audio.inputMomentname, transport (built_in, bluetooth, usb, display, virtual, or other), and reason (selected, system_default, or lid_closed). The input the audio comes from, from this moment on.
tutorial.startedMomentname (voice_introduction), and prompt?: the text the user was asked to read. Audio from here on is part of an app tutorial, not ordinary conversation.
tutorial.endedMomentname, and completed?: whether the tutorial ran to its end. A start without an end means capture stopped during the tutorial.

A recording of several people can carry one speaker.active signal per stretch of speech. To describe a stream that is a single person's own channel, send one signal that covers the whole stream, from start_ms 0 to the stream's end_ms, with that person as the participant.

Plum for Mac posts one audio.input signal at the start of each stream it records. It starts a new stream whenever its input changes, including when it replaces a MacBook's built-in microphone, which records silence while the lid is closed, with reason lid_closed. Its setup's voice introduction, where the user reads a sentence aloud, is bracketed by tutorial.started and tutorial.ended in the stream it records.

TypeScript clients can import the payload schemas, a reader that returns a typed value when a stored signal matches a known type and a generic { type, payload } otherwise, and builders for known types from @optima/contracts/signals.

Read signals

GET /audio/streams/{id}/signals lists a stream's signals ordered by start_ms, then id:

Query parameterDescription
typeOnly signals of exactly this type.
from_msOnly signals that end at or after this time. A moment ends where it starts.
to_msOnly signals that start at or before this time.
limitPage size, 1 to 100. Defaults to 50.
cursorOpaque cursor from the previous page. Keep the same filters.

The response has the usual { items, cursor } shape. from_ms and to_ms together return the signals that overlap a stretch of the recording:

curl 'https://api.getoptima.com/audio/streams/STREAM_ID/signals?type=speaker.active&from_ms=0&to_ms=60000' \
  -H 'Authorization: Bearer YOUR_ACCESS_TOKEN' \
  -H 'X-API-Version: 1'

A stream that is not yours, or that belongs to a deleted source, returns 404. Signals produce no account events. The MCP tool list_stream_signals reads them for agents. See attach signals to a stream.

Retries and errors

If a response is lost, repeat the same request with the same refs, metadata, and bytes. Every step is safe to repeat this way, including after the chunk is confirmed:

Repeated stepResult
Create a source or streamReturns the existing record
End a stream with the same end_msReturns the stream
Post signals with refs already stored with identical contentReturns the stored signals; nothing is added. A signal whose payload expired matches on its type and timing alone.
Create an upload intentReturns the same upload_path
Upload bytesReturns the stored object's tag without replacing the bytes
Confirm custodyReturns the original receipt, including its first received_at

Stored audio is immutable. Once bytes are stored for a chunk, a later upload with the same Content-Length returns the existing object's tag. Optima does not compare the new body with the stored bytes and never replaces them.

ProblemResult
Metadata conflicts with an earlier request for the same ref or stream positionHTTP 409 CONFLICT
A stream's external_ref is already used by another stream of the sourceHTTP 409 CONFLICT
A chunk extends past its stream's recorded endHTTP 409 CONFLICT
A signal ref is already stored with different content; the request stores nothingHTTP 409 CONFLICT
A signal has a malformed type, an end_ms before start_ms, a payload over 8192 bytes, a repeated ref, or the request has more than 500 signalsHTTP 422 VALIDATION_ERROR
Content-Length is missing or differs from size_bytesHTTP 422 VALIDATION_ERROR
Uploaded bytes do not match the intent's SHA-256 digestHTTP 422 VALIDATION_ERROR
Two uploads for the same chunk raceOne can return HTTP 409 CONFLICT. Repeat the upload or confirm.
Confirmation is requested before a complete uploadHTTP 409 CONFLICT

See API conventions for the error envelope.

Check processing

GET /audio/{chunk_id} reports the chunk's status:

StatusMeaning
pending_uploadThe intent exists, but custody is not confirmed
waitingCustody is confirmed and the chunk is queued for transcription
processingTranscription is in progress
readyProcessing finished
failedProcessing failed. The confirmed bytes remain stored.

GET /audio lists confirmed audio only. Pass an absolute from and to window, or unknown_time=true for audio without a capture time. Each confirmed record includes file: id, mime, size_bytes, sha256, and url, a signed storage link that downloads the bytes directly until url_expires_at (15 minutes). Request it without API headers and read the record again for a fresh link. file is null until custody is confirmed.

Audio that the workspace's retention removed is not listed, and GET /audio/{chunk_id} for it returns 410 CONTENT_EXPIRED (or CONTENT_DELETED) with removed_at. See Removed content.

Published text appears in turns. Hardware devices use the same sequence under /device; see Devices. Source activity shows queued and transcribing counts per source. See the API reference for exact request and response schemas.

PCM16 WAV uploads preserve uncompressed samples and use more bandwidth than Ogg Opus. The Omi bridge uses WAV for its incoming raw PCM. Both formats follow the same intent, content upload, and confirmation flow.

On this page