Position of this clip in the language's spoken sequence, starting at 1.
Milliseconds from source finalisation to the audio URL reaching listeners.
Additional milliseconds this clip waits behind audio already queued.
Total milliseconds the spoken audio starts behind the source: pipeline plus backlog.
Playback duration of the synthesised clip.
Source speech this clip covers, or null for the first clip or after a long gap.
audioMs / sourceMs, or null when sourceMs is unknown. Above 1 means drift.
One measured spoken segment.