Clera
Research Engineer, Synthetic Data
- Compensation
- $150K–$250K / yr
- From job posting
- Location
- Singapore, SG
- Arrangement
- On-site
- Employment type
- Full-time
- Posted
- 26 September 2026 (4 days ago)
Checked 3 days agoApplications go to the employer, never to RoleSprint
About this role
ABOUT THE ROLE
This is a Research Engineer role focused on synthetic data, sitting within a roughly 15-person engineering team of Olympiad medalists and published researchers. You will build the pipelines that turn domain-specific workflows into scalable, high-quality training tasks for AI agents, directly shaping what models learn and how well they perform.
WHAT YOU'LL DO
- Build end-to-end synthetic data pipelines that transform domain-specific workflows into realistic, structured, and challenging training tasks.
- Collaborate with subject-matter experts to create synthetic tasks for AI agents across professional and technical domains.
- Design task generation methods that produce diverse, realistic, and learnable outputs at scale.
- Build tooling to mutate, validate, and continuously improve synthetic tasks.
- Analyze model and agent performance on synthetic tasks to identify what they teach and where they break down.
- Develop metrics to quantify synthetic task diversity, realism, learnability, and overall quality.
WHAT WE'RE LOOKING FOR
- 2 to 4 years of experience in software engineering, machine learning engineering, or AI research, with hands-on work building data pipelines, ML infrastructure, or synthetic data systems.
- Proficiency in Python and experience developing in Linux environments using containerization tools such as Docker.
- Demonstrated experience applying synthetic data research methods to build end-to-end data generation pipelines for AI/ML applications.
- Strong understanding of synthetic data quality criteria and evaluation metrics, including diversity, realism, and learnability, as well as their inherent limitations.
- Experience designing, implementing, or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models.
- Track record of independently owning and delivering technical projects end-to-end with minimal predefined requirements.
- Experience building automated systems to generate, validate, mutate, or process structured datasets at scale.
- Sharp eye for edge cases, inconsistencies, and quality issues in synthetic or algorithmically generated data.
- Familiarity with reinforcement learning training paradigms, agentic AI workflows, or LLM post-training pipelines is a plus.
- Comfortable operating in unstructured, early-stage environments and reasoning from first principles.
- Strong communication skills for asynchronous, cross-timezone collaboration.
COMPENSATION & BENEFITS
Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available.
LOCATION
On-site in Singapore.
Work location
- Singapore, SG
Related jobs
Ready to make a decision?
This role is either worth your time or it isn’t.
Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.
Nothing is submitted automatically. You choose what happens next.
About this listing
Published on Ashby under the board identifier Clera, which is the name the employer’s own job board carries. RoleSprint has not verified the company’s registered or trading name, so it is shown exactly as published rather than tidied up.
RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.
Published 26 September 2026, last checked 3 days ago. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.