- CRAWL + EXTRACT
- DEDUP
- +6 more
Closing the data gap for RSI.
Horizon partners with experts in low-level coding infrastructure domains on the critical path to increasing the rate of recursive self-improvement.Kernel, compiler, inference, and systems engineers produce the rare training signal frontier models need to improve the infrastructure beneath their own intelligence.The next gains in model capability depend on systems work that never appears cleanly in public code. Horizon turns that expert work into executable learning environments.Frontier models cannot automate improvement until they master the low-level capabilities that constrain it. Horizon sources those capabilities from the engineers who build them.From pre-training algorithms and CUDA kernels to inference engines and agent harnesses, Horizon captures expert work across the entire recursive self-improvement loop.Public repositories show finished code, not the decisions that make difficult infrastructure work succeed. Horizon captures those trajectories as verifiable training data.We build the low-level coding environments frontier models need to take on more of the research and engineering required to improve themselves.By focusing expert data production on RSI-critical infrastructure, Horizon increases the rate at which frontier models acquire their next generation of capabilities.Rare systems expertise becomes useful training data only when it is executable, inspectable, and rewardable. Horizon makes that transformation possible.Horizon concentrates expert trajectories, RL environments, and verifier design on the engineering hotpath for recursive self-improvement.
OPERATIONAL SCALE
EXPLORE OUR DATA
RL environments across 62 key model capabilities.RL environments for the capabilities that compound model improvement.RL environments across 62 key model capabilities, mapped to executable training targets.A data catalog spanning the model engineering stack.Executable environments, expert trajectories, and verifiable reward by capability.
Get in touch
12 stages · 62 model capabilities
- DISTRIB TRAINING
- PARALLELISM
- +5 more
- DEVICE KERNELS
- KERNEL DSLS
- +3 more
- SFT
- PEFT + ADAPTERS
- +3 more
- LLM RL
- CLASSIC RL
- +4 more
- SERVING
- QUANTIZATION
- +5 more
- EVAL HARNESSES
- CODE + SWE EVALS
- +3 more
- CODING AGENTS
- TOOL PROTOCOLS
- +4 more
- PROMPT OPT
- PROMPT TESTING
- +2 more
- SAES
- MODEL INTERNALS
- +1 more
- RED TEAMING
- GUARDRAILS
- EXPERIMENT TRACKING
- JOB ORCHESTRATION
- +2 more
Explore all 62 capabilitiesFull environment taxonomy
Full model capability map
62 engineering capabilitiesData acquisition & curation
- Web-scale crawl & extraction
- Dedup & near-dedup
- Quality filtering & classifiers
- Decontamination & PII
- Tokenizer training
- Synthetic data generation
- Data mixing & curriculum
- Dataset IO & streaming
Pre-training
- Distributed training frameworks
- Parallelism primitives
- Minimal reference trainers
- Optimizers & schedules
- Numerics & precision
- Architecture research
- Checkpointing & fault tolerance
Kernels & compilers
- Hand-written device kernels
- Kernel DSLs & tile languages
- Compilers & graph codegen
- Autotuning & kernel benchmarks
- Collective communication
Post-training & alignment
- SFT frameworks
- PEFT & adapters
- Preference optimization
- Reward modeling
- Distillation & compression
RL & environments
- LLM RL frameworks
- Classic RL libraries
- Agentic environments
- Verifiers & graders
- Rollout & sampling infra
- Sandboxing & code execution
Inference & serving
- Serving engines
- Quantization
- KV cache & memory
- Speculative & parallel decoding
- Constrained & structured output
- Routing & scheduling
- Edge & on-device
Evaluation & benchmarking
- General eval harnesses
- Code & SWE benchmarks
- Agentic & tool-use evals
- Contamination & leakage
- Eval statistics & infra
Agent harness design
- Coding agent scaffolds
- Tool protocols & registries
- Context & memory management
- Multi-agent orchestration
- Computer & browser use
- Agent observability
Prompting & context engineering
- Programmatic prompt optimization
- Prompt testing & management
- Chat templates & serialization
- Skill & instruction packs
Interpretability & model analysis
- SAEs & dictionary learning
- Model internals tooling
- Steering & model editing
Safety, red-teaming & guardrails
- Red-teaming & jailbreak
- Guardrails & filtering
MLOps, infra & reproducibility
- Experiment tracking
- Config & job orchestration
- Model formats & registries
- Cluster & accelerator infra