Exercises are prepared for evaluation; we do not use candidates’ submissions in production. Depending on the role, progress might be a tested change, a well-designed experiment, or a diagnosis supported by evidence. Examples include:
- Inference & kernels
- Investigate a performance trace and test an optimization against correctness and representative workloads.
- Agents & multi-agents
- Diagnose a failed workflow and demonstrate a reliable recovery in a prepared environment.
- Desktop worker
- Improve an interaction involving permissions, background work, or state recovery.
- Streaming vision
- Investigate a screen sequence and evaluate a change within a latency budget.
- Serving & fleet
- Explore an incident and test a scheduling or recovery improvement.
- Voice & hearing
- Trace a turn-taking failure and evaluate conversational behavior and latency.
- Post-training & distillation
- Examine an evaluation result and design a small experiment to test its explanation.