RL Environments & AI Evaluation Infrastructure
Bespokelabs AI · Micro1 · Handshake AI
Designed realistic RL environments, task specifications, reward signals, verifiers, hidden tests, and benchmark tasks for AI coding agents and model evaluation workflows.
- Python
- PyTorch
- JAX
- Hugging Face
- verifiers
- Docker
- pytest
- CI
- benchmark harnesses
Outcomes
- Authored realistic multi-step coding-agent repair tasks
- Built deterministic verification and golden-reference checks
- Analyzed rollouts for failure modes and reward hacking
- Created evaluation criteria for reproducibility, correctness, and robustness