Hoid
ActiveAI-nativeAgentic AI compiler — makes AI models faster & cheaper on any hardware
What it does
Hoid builds software that optimizes AI inference for any model on any hardware. Today, it takes expert engineers weeks to optimize a single model. Oftentimes, by the time they finish, the next hot model is already out. Hoid's agentic AI compiler optimizes a new model in days instead of weeks and automatically beats expert-optimized engines like vLLM and SGLang by 30%. In just one month, Hoid has scaled to a $196K contracted run rate with live customers, and additional advanced pilots are underway with an audio model lab, multiple inference providers, and a large, global autonomous driving company. Hoid was co-founded by a team of AI performance engineers. CEO Momcilo worked on ultra-low-latency data ingestion at Databricks and multi-GPU training optimization at Microsoft. CTO Pavle graduated from Cambridge and, together with CRO Vladimir, optimized the world's fastest video generation model at Tenstorrent.
Its site says now · captured 2026-09-19
Hoid — Run any model on any hardware
The software stack for AI inference. Run any model on any hardware — faster and cheaper.
Next to it in AI infra and compute · AI compiler for hardware optimization
- LuminalY Combinator S25
Making AI run fast on any hardware.
- Texel.aiY Combinator W23
Run AI pipelines 10x faster
- BelvedirY Combinator S26
The autonomous AI model factory
- Experiential LabsY Combinator S26
Open source AI gateway that turns your traffic into better models
- Harbor AIa16z speedrun SR007
The operating layer that keeps GPU fleets productive.
- TracerY Combinator S26
Combining open-source AI models for better answers at lower cost
- OrchestraY Combinator S26
Self-optimizing neocloud that cuts LLM costs by 100x
- RightNowY Combinator F26
Co-designing the fastest intelligence