ASYNC SPEECH-TO-TEXT API · PODCASTS, VIDEOS, PUBLIC MEDIA
Transcribe an hour for 2¢.
Get it back in under a minute.
Turn hours of podcasts, videos and public media into accurate, timestamped transcripts in minutes. $0.02 per audio hour.
$0.02/hour · <60 s p95 for a 1-hour file · timestamps + SRT/VTT · no minimums
$ curl https://api.lopta.dev/v1/transcriptions \
-H "Authorization: Bearer $LOPTA_KEY" \
-F file=@episode-142.mp3 \
-F language=en \
-F webhook=https://your.app/hooks/lopta
→ 202 { "id": "tr_…", "status": "queued" }
→ webhook { "status": "done", "text": "Welcome back…", "srt_url": "…" }
One hour in.
Transcript out.
An hour of media became a transcript in seconds, and cost a few cents.
- 00:00:04
- Welcome back to the show. This is episode one hundred forty-two.
- 00:00:08
- Today we're talking about how local radio stations archive forty years of tape.
- 00:00:15
- And the hardest part isn't storing it. It's finding anything once it's stored.
Batch pricing.
Without the batch queue.
Batch APIs are cheap because they make you wait. Lopta is built to give recorded media batch economics without batch turnaround.
Public list prices as of September 2026, base transcription only, no add-ons. Deepgram converted from $0.0043/min. Groq bills a 10-second minimum per request; Lopta bills per second.
Built differently.
So it can cost differently.
Most AI infrastructure provisions centralized compute to meet demand. Recorded media gives us another option.
Lopta is designed to split asynchronous workloads into independent jobs and route them to efficient available compute before stitching the result back together.
Our network is built around edge devices.
Provision compute to meet demand.
Route jobs to available compute.
Provision less. Schedule more.
HALF THE PRICE IS WORTHLESS IF THE WORDS ARE WRONG
Every language has to earn its way in.
We only open a language once it matches our reference API across public, reproducible benchmarks.
Full methodology on request →| ENGLISH · WER ↓ | LOPTA | GROQ · WHISPER TURBO |
|---|---|---|
| LibriSpeech read speech | 1.76 % | 1.85 % |
| TED-LIUM public talks | 3.46 % | 3.66 % |
| Earnings-22 accented, domain-heavy | 9.80 % | 10.30 % |
| AMI multi-speaker, far-field | 7.81 % | 14.81 % |
Three calls. No SDK required.
-
01
Send the file
Upload it or pass a URL, and name the language up front — no transcript silently returned in the wrong one.
-
02
Split, process, stitch
We split long recordings at natural boundaries, process chunks asynchronously in parallel, then rebuild the transcript with timestamps.
-
03
Webhook or poll
Get the result pushed to your endpoint, or poll the job within the hour. Plain text, timestamped segments, SRT and VTT in the same response.
One language at a time, and only when it's good.
- English LIVE
- French IN EVALUATION
We don't charge for real-time latency you don't need.
Live voice needs milliseconds. Recorded media doesn't. It needs accurate results in minutes, reliably.
Lopta is optimized specifically for that trade-off.
Questions
Data Processing page →- What audio is Lopta built for?
- Content that is already public or intended to be public: podcasts, videos, broadcasts, talks and archives.
- Lopta is not currently designed for confidential meetings, HR calls, medical consultations, legal recordings or private corporate audio.
- What do I get for free?
- $1 of credit when you create your account — 50 hours of audio. No card required until you want more.
- Is there a streaming endpoint?
- No, by design. Lopta is for recorded media. For live captions or voice agents, use a streaming provider.
- What are the limits?
- Audio files only (MP3, M4A, WAV, FLAC, OGG, Opus) — extract the audio track before sending video. Up to 500 MB by upload or 1 GB by URL, and up to 5 hours per file. 10 concurrent jobs by default; additional jobs are queued, not rejected.
- Where is my audio processed?
- Jobs are processed across Lopta's distributed compute network, which includes edge devices — by using Lopta you accept processing on consumer devices. The full list of subprocessors is on the Data Processing page.
- How long is it retained?
- Audio is deleted as soon as processing ends, whether the job succeeds or fails. Transcripts are kept for 1 hour so you can fetch them, then deleted. We never train on your data.
$1 = 50 hours
of audio.
It's on us.
Request received.
You'll hear from us shortly with your API key.