Greenthread
UnknownAI-nativeThe GreenThread inference engine allows you to multiplex many LLM's and SLM's on the same GPU, swapping them subsecond.
Next to it in AI infra and compute · GPU inference multiplexing for LLMs
- Cumulus LabsY Combinator W26
The Fastest Multimodal Inference OS
- OvershootY Combinator W26
AI Infra for real-time vision applications
- generalcomputePlug and Play PnP 2026
General Compute provides an ASIC-based cloud inference platform with an OpenAI-compatible API for deploying and serving AI models.
- OpenRelayY Combinator S26
Distributed, hardware-agnostic AI inference
- TracerY Combinator S26
Combining open-source AI models for better answers at lower cost
- OrchestraY Combinator S26
Self-optimizing neocloud that cuts LLM costs by 100x
- LuminalY Combinator S25
Making AI run fast on any hardware.
- RightNowY Combinator F26
Co-designing the fastest intelligence