Speech-to-Text in 60+ languages at 130× real-time and multilingual Text-to-Speech. Same cluster, same transparent pricing, one unified balance.
import OpenAI from "openai"; const orchard = new OpenAI({ apiKey: process.env.ORCHARD_API_KEY, baseURL: "https://api.orchardrun.com/v1",}); // 1. Transcribe · 60+ languages · 130× real-timeconst transcript = await orchard.audio.transcriptions.create({ file }); // 2. Synthesize · 17 languages · latency under 2 sconst audio = await orchard.audio.speech.create({ input: "Welcome back, your order is on its way.", voice: "claribel",}); // 3. Diarize · label who spoke whenconst diarized = await orchard.audio.transcriptions.create({ file, diarize: true });Drop into any TypeScript / Node app · Python SDK identical
“Voice synthesis at production scale. One API, one balance, seventeen languages.”
Audio rendered with Orchard · pre-cached, instant playback
Transcription + synthesis. Same API, same billing, same dedicated infrastructure.
Press Cmd+Shift+8, speak, paste at cursor. Works in Cursor, Claude Code, Copilot — and any editor.
Install on MarketplaceSix real workflows where Orchard replaces expensive transcription and synthesis APIs, or fragmented service stacks — all on one API.
Transcribe WhatsApp, Telegram or live call audio and feed the context to your LLM. Low latency, controlled cost per minute.
Transcribe podcasts, meetings, interviews or thousands of files a day. No rate limits on paid plans, no per-file caps.
Turn customer calls into text + automatic insights for your team. Speaker diarization, ideal for call centers and QA.
Natural multilingual voiceovers for videos, ads and automations — one consistent brand voice across hundreds of assets.
Audio → Transcript → Summary → Action with your LLM of choice. Webhooks, retries and batching native to the API.
Build your own conversational assistant or branded voice agent. Text-to-speech in the voice you define, in the language you need.
60× real-time average sustained. 1 hour of audio in under 1 minute.
$0.00042/min on Pro plan. Simple plans, no surprise costs.
Industry-standard API. Existing SDKs work without changes — migrate in minutes.
Builders, indie hackers and audio-first teams across Latin America, the US and Europe ship faster on Orchard's pay-as-you-go speech stack.
Start free · Upgrade as your volume grows
500 min on signup
≈ 500K chars TTS
Try all 3 products. No card required.
1,500 min/month
≈ 1.5M chars TTS
Coffee-money tier. STT + TTS share the balance.
15,000 min/month
≈ 15M chars TTS
Bots and small SaaS. All 3 products on shared balance.
60,000 min/month
≈ 60M chars TTS
Production volume. All 3 products on shared balance.
Optional diarization · Custom SLA · Dedicated capacity
Our API mirrors the OpenAI Whisper request/response format, so most SDKs work without code changes. Point them at our endpoint and swap the key. Median migration is under an hour.
500 free minutes on signup. No card. Same balance across STT and TTS. Pay once, use everything.
We founded Orchard with one purpose: to reshape the Voice Infrastructure industry. As heavy consumers ourselves, we kept hitting the same gaps in the market — exactly where we decided to differentiate: price, volume and concurrency. That's why we built our core verticals: STT and TTS.
Our strongest surface today is STT batch — and we're going for the global #1 spot. We back it up with three hard numbers: the cheapest minute on the market, a WER competitive with the best engines in the segment, and an RTF that sustains high volume and massive concurrency without throttling. That combination of quality, speed and price isn't on offer anywhere else.
In parallel, our TTS is consolidating as the default base for voice agents, voice assistants and conversational products — a segment growing double digits as every product turns voice-first.
Latin American speech is the bet we're most excited about for what's next. Generic models still fall apart on real audio here — street noise, phone-quality calls, fast colloquial speech across a dozen accents. Where we're investing heavily is the pipeline: fine-tuning for the accents and prosody of the region so the transcript holds up on the audio our customers actually process, not just clean studio recordings.

Processing audio at scale breaks you: pipelines that don't scale, latencies that kill the product, costs that crush you exactly when you grow. The next wave of software is going to be voice-first and agentic. And we want to be the infrastructure it stands on, not the bottleneck.

We're aiming to be the world #1 in STT batch. Today we combine the lowest cost on the market with quality on par with the leaders, and we're investing heavily in Latin American speech, where capturing real accents is still an unsolved problem.
Encryption, audit logs, and multi-region replication are table stakes. Data privacy and zero-retention on your audio — that's our line.
DPA available, full data deletion on request
California consumer rights honored end-to-end
All stored audio and transcripts encrypted
Modern cipher suites only, no downgrade
Database tier replicated with point-in-time recovery
Role-based access control across organization seats
Synthetic-audio disclosure + consent framework
In progress — targeting Q4 2026