Abundant
ActiveAI-nativeAgent simulation and RL for researchers
What it does
Hello! 👋 We are a team of former ML engineers, founders, roboticists and ops leads who obsess about data and its impact on safe, reliable AI. We specialize in creating environments and datasets for RL by leveraging our experience in simulation and model training. By the numbers: • Powering 3 of the top 6 global AI labs and multiple Fortune 500 enterprises • Billions of training tokens generated each month, 2x month over month • Exclusive, global network of over 500 domain experts We believe humans are inherently creative, and thrive by pushing the frontier. We are working towards an abundant future--one where everyone has access to infinite intelligence, services and goods. Based in San Francisco, CA. We enjoy good food and good company. -- more info below -- Abundant is building the NVIDIA of training data. AI models rely on two fundamental ingredients: compute and data. NVIDIA, the leader in compute, has a peak market cap of $5T and generated $130B in revenue last year as the need for scaling compute has exploded. We believe the need to scale data is just beginning, as we move beyond SFT and human supervision to RL and Learning from Experience. Our founding team consists of second-time founders, ML engineers and data leads from Waymo, Google, Meta and AWS. Our team has previously collaborated with DeepMind to classify hate speech in YouTube videos, trained SOTA models for self-driving, and scaled data pipelines with thousands of human annotators. Our pioneering work in human computation, synthetic data, imitation learning and RL give us a solid advantage in delivering results to our customers. Why now? Training data is more important and more scarce than ever before. Scaling laws dictate that linear improvement in model performance demands an exponential increase in training data. But there is only one World Wide Web and most of it has already been trained on. The next advances will require new, diverse, and high-quality dat
Its site says now · captured 2026-08-23
abundant
Environments and datasets for RL. Powering the world
Next to it in Data for AI · RL simulation and synthetic data generation
- NeosyncY Combinator S22
Neosync is an open-source anonymization and synthetic data platform.
- Synthetic SocietyY Combinator S25
Synthetic Users to Simulate Real Users
- BesampleTechstars TS 2024
Besample | Research beyond the West
- DatacurveY Combinator W24
Frontier coding data for training and evaluating LLMs
- PerfectBit, Inc.Y Combinator X26
Correct by construction AI training data
- Olam LabsY Combinator S26
Building multi-agent simulations for model evals and training.
- QokedasY Combinator F26
Data for AI Science
- Sepal AIY Combinator S24
Data Development for Advanced AI