Build code-understanding and review workflows across repository context, pull request diffs, project instructions and custom checks.
Improve independent review and verification passes, prioritisation and explanations so useful bugs surface with fewer false positives.
Develop evaluation datasets from public, synthetic or explicitly authorised examples; measure precision, coverage, latency and cost without using customer content to train models.
Investigate prompt injection, tool behaviour and failure cases in sandboxed agents, and work with product engineers to ship improvements.
What you’ll bring
Experience shipping LLM applications, agent workflows or developer tools, with strong software engineering fundamentals.
Ability to read unfamiliar code, reason about bugs and design repeatable evaluations rather than rely on a compelling demo.
Comfort working with TypeScript or Python, model APIs, Git and automated tests. Static analysis or security experience is useful.
Apply for this role
Tell us about yourself and a project you contributed to. A few thoughtful paragraphs are enough.
This form prepares an email. Your application is sent only when you send it from your email app.