ASYNC SPEECH-TO-TEXT API · PODCASTS, VIDEOS, PUBLIC MEDIA

Transcribe an hour for 2¢.
Get it back in under a minute.

Turn hours of podcasts, videos and public media into accurate, timestamped transcripts in minutes. $0.02 per audio hour.

$0.02/hour · <60 s p95 for a 1-hour file · timestamps + SRT/VTT · no minimums

ONE REQUEST
episode-142.mp31h 02m · English
Lopta<60 s p95
transcript.txtsegments.json · timestampsepisode-142.srtepisode-142.vtt

$ curl https://api.lopta.dev/v1/transcriptions \
    -H "Authorization: Bearer $LOPTA_KEY" \
    -F file=@episode-142.mp3 \
    -F language=en \
    -F webhook=https://your.app/hooks/lopta
→ 202  { "id": "tr_…", "status": "queued" }
→ webhook  { "status": "done", "text": "Welcome back…", "srt_url": "…" }

One hour in.
Transcript out.

An hour of media became a transcript in seconds, and cost a few cents.

1h 02mAUDIO IN
<60 sPROCESSED
$0.021BILLED
INPUTepisode-142.mp3English
OUTPUTTXTTimestampsSRTVTT
EXCERPT · episode-142.srt
00:00:04
Welcome back to the show. This is episode one hundred forty-two.
00:00:08
Today we're talking about how local radio stations archive forty years of tape.
00:00:15
And the hardest part isn't storing it. It's finding anything once it's stored.

Batch pricing.
Without the batch queue.

Batch APIs are cheap because they make you wait. Lopta is built to give recorded media batch economics without batch turnaround.

AUDIO PER MONTH
PROVIDER / MODELPER HOURTURNAROUNDMONTHLY COST AT 10,000 H
MONTHLY COST AT 10,000 H

Public list prices as of September 2026, base transcription only, no add-ons. Deepgram converted from $0.0043/min. Groq bills a 10-second minimum per request; Lopta bills per second.

Built differently.
So it can cost differently.

Most AI infrastructure provisions centralized compute to meet demand. Recorded media gives us another option.

Lopta is designed to split asynchronous workloads into independent jobs and route them to efficient available compute before stitching the result back together.

Our network is built around edge devices.

PROVISIONED COMPUTE

Provision compute to meet demand.

Audio
Provisioned centralized capacity
Transcript
SCHEDULED COMPUTE

Route jobs to available compute.

Audio
Split into independent jobs
Scheduler
Edge devices
Stitch
Transcript

Provision less. Schedule more.

HALF THE PRICE IS WORTHLESS IF THE WORDS ARE WRONG

Every language has to earn its way in.

We only open a language once it matches our reference API across public, reproducible benchmarks.

Full methodology on request →
ENGLISH · WER ↓ LOPTA GROQ · WHISPER TURBO
LibriSpeech read speech 1.76 % 1.85 %
TED-LIUM public talks 3.46 % 3.66 %
Earnings-22 accented, domain-heavy 9.80 % 10.30 %
AMI multi-speaker, far-field 7.81 % 14.81 %
Reference: Groq whisper-large-v3-turbo. Up to 250 utterances ≥ 5 s per test set (seed 42), identical audio and text normalization for both systems. Parity = within 1.0 WER point of the reference. Full-recording results coming.

Three calls. No SDK required.

  1. 01

    Send the file

    Upload it or pass a URL, and name the language up front — no transcript silently returned in the wrong one.

  2. 02

    Split, process, stitch

    We split long recordings at natural boundaries, process chunks asynchronously in parallel, then rebuild the transcript with timestamps.

  3. 03

    Webhook or poll

    Get the result pushed to your endpoint, or poll the job within the hour. Plain text, timestamped segments, SRT and VTT in the same response.

LANGUAGES

One language at a time, and only when it's good.

  • English LIVE
  • French IN EVALUATION
PRODUCT PHILOSOPHY

We don't charge for real-time latency you don't need.

Live voice needs milliseconds. Recorded media doesn't. It needs accurate results in minutes, reliably.

Lopta is optimized specifically for that trade-off.

What audio is Lopta built for?
Content that is already public or intended to be public: podcasts, videos, broadcasts, talks and archives.
Lopta is not currently designed for confidential meetings, HR calls, medical consultations, legal recordings or private corporate audio.
What do I get for free?
$1 of credit when you create your account — 50 hours of audio. No card required until you want more.
Is there a streaming endpoint?
No, by design. Lopta is for recorded media. For live captions or voice agents, use a streaming provider.
What are the limits?
Audio files only (MP3, M4A, WAV, FLAC, OGG, Opus) — extract the audio track before sending video. Up to 500 MB by upload or 1 GB by URL, and up to 5 hours per file. 10 concurrent jobs by default; additional jobs are queued, not rejected.
Where is my audio processed?
Jobs are processed across Lopta's distributed compute network, which includes edge devices — by using Lopta you accept processing on consumer devices. The full list of subprocessors is on the Data Processing page.
How long is it retained?
Audio is deleted as soon as processing ends, whether the job succeeds or fails. Transcripts are kept for 1 hour so you can fetch them, then deleted. We never train on your data.

$1 = 50 hours
of audio.

It's on us.

We reply personally. No newsletter. Privacy