Full-Stack Software Engineer, Reinforcement Learning
Software Engineering
About the Role
You'll be a Full-Stack Software Engineer building the product surfaces, backend systems, and internal tools that power HUD's RL data engine. You'll own end-to-end product experiences—from frontend dashboards to backend services—that enable frontier labs and external partners to create, evaluate, and iterate on reinforcement learning training data. You don't need to be an ML researcher, but you'll work closely with research engineers and vendors to translate complex, ambiguous needs into polished, intuitive products.
What You'll Do
- Develop product-facing tools for browsing environments, inspecting trajectories, reviewing task quality, and understanding model behavior.
- Build vendor-facing workflows that make it easy for external partners to create, submit, test, and iterate on RL environments and training data.
- Create dashboards and observability tools that surface environment quality, eval results, data collection progress, and pipeline health.
- Design backend services and APIs connecting task authoring, data collection, evaluation, QA/QC, and RL training infrastructure.
- Partner with research, operations, and GTM teams to ship well-designed systems quickly without waiting for perfect specs.
What We're Looking For
- Strong full-stack software engineering fundamentals with proficiency in Python and a modern web stack (React, TypeScript, Next.js, or similar).
- Experience owning user-facing or internal products end-to-end.
- Good product taste and ability to build intuitive tools for both technical and non-technical users.
- Comfort with cloud infrastructure, Docker, CI/CD, observability, and production debugging.
- High agency—you identify what needs to exist, build it, and improve it independently.
- Strong communication skills across research, engineering, operations, vendors, and founders.
- Experience with data collection, labeling, eval, or research tooling platforms is a plus.
- Background building dashboards, review workflows, observability tools, or debugging interfaces for complex systems is valuable.
Compensation & Benefits
Competitive compensation, 100% covered medical/dental/vision (US employees), lunch and dinner in-office, company-wide holiday break (Christmas Eve–New Year's Day) plus PTO, Equinox membership, 401k, commuter benefits (US employees), and unlimited access to ChatGPT, Claude, and Cursor tokens.
Location
Offices in San Francisco or Singapore with openness to remote candidates who can maintain 70–80% time zone overlap with either location. Visa sponsorship and relocation support provided for strong full-time candidates to the US or Singapore.