Oruk
ActiveAI-nativeSpeech foundation models that understand people: words, tone of voice, emphasis, intent, and emotion.
What it does
As AI tools become mainstream and robots enter the workforce, conversation is replacing the text box as the only interface fast enough to keep up with critical demands. The voice layer behind every agent, coding session, robot command, and phone call needs the full context: words, tone of voice, emphasis, intent, and emotion. When someone says "I'm fine," they could be reassuring you or asking you to leave them alone. The words are the same. The meaning is in the voice. We are speech, AI, and linguistics researchers from Stanford, Cambridge, and Berkeley, building Oruk to train speech foundation models that actually understand people. We've already set new records across speech understanding benchmarks, deployed to hundreds of thousands of users, and assembled the world's largest speech perception dataset.
Its site says now · captured 2026-09-19
oruk — Speech API for transcripts, emotion, and style
Oruk builds speech models and an API for English transcription, emotion scores, and speaking style, with Python and TypeScript SDKs and reproducible research.
Next to it in AI infra and compute · Speech foundation models
- Induction LabsY Combinator S26
Building intellectually curious AI
- Miso LabsY Combinator X26
Emotive foundation voice models
- Kalpa LabsY Combinator F25
Scaling Generalist Speech models
- Maya ResearchSouth Park Commons
Models that understand how it's said
- WisentEntrepreneur First EF London 2025
World's best AI.
- Zorbea16z speedrun SR007
Company-Owned AI models that beat the frontier
- AxionOrbital SpaceY Combinator W26
Foundation models for 24/7 Earth Observation
- Iron GridY Combinator S25
AI-Native Data Center Engineering