Oruk

ActiveAI-native

Speech foundation models that understand people: words, tone of voice, emphasis, intent, and emotion.

What it does

As AI tools become mainstream and robots enter the workforce, conversation is replacing the text box as the only interface fast enough to keep up with critical demands. The voice layer behind every agent, coding session, robot command, and phone call needs the full context: words, tone of voice, emphasis, intent, and emotion. When someone says "I'm fine," they could be reassuring you or asking you to leave them alone. The words are the same. The meaning is in the voice. We are speech, AI, and linguistics researchers from Stanford, Cambridge, and Berkeley, building Oruk to train speech foundation models that actually understand people. We've already set new records across speech understanding benchmarks, deployed to hundreds of thousands of users, and assembled the world's largest speech perception dataset.

Its site says now · captured 2026-09-19

oruk — Speech API for transcripts, emotion, and style

Oruk builds speech models and an API for English transcription, emotion scores, and speaking style, with Python and TypeScript SDKs and reproducible research.

Next to it in AI infra and compute · Speech foundation models

  • Induction LabsY Combinator S26

    Building intellectually curious AI

  • Miso LabsY Combinator X26

    Emotive foundation voice models

  • Kalpa LabsY Combinator F25

    Scaling Generalist Speech models

  • Maya ResearchSouth Park Commons

    Models that understand how it's said

  • WisentEntrepreneur First EF London 2025

    World's best AI.

  • Zorbea16z speedrun SR007

    Company-Owned AI models that beat the frontier

  • AxionOrbital SpaceY Combinator W26

    Foundation models for 24/7 Earth Observation

  • Iron GridY Combinator S25

    AI-Native Data Center Engineering