Skip to content

Jobgether

Software Engineer, AI Systems

Location
US
Arrangement
Remote
Employment type
Full-time
Posted
30 September 2026 (today)

Checked todayApplications go to the employer, never to RoleSprint

About this role

Accountabilities • Build and operate production LLM pipelines coordinating model calls, tool calls, graph queries, retrieval, quality gates, and specialist-agent handoffs.

• Extend agent orchestration systems that move incidents from evidence collection through analysis, review, and organizational learning.

• Design grounding and retrieval strategies using graph traversal, vector search, and hybrid retrieval to provide models with relevant evidence and organizational knowledge.

• Develop evaluation datasets, scoring systems, regression suites, model comparisons, LLM-as-judge workflows, and human-review loops for extraction and reasoning tasks.

• Implement production AI operations capabilities including tracing, tool-call auditing, cost and latency monitoring, failure handling, and quality dashboards.

• Identify and address hallucinations, agent loops, silent model drift, regressions, and other production-quality issues before they affect customers.

• Evaluate and select models across providers based on quality, latency, cost, context capabilities, and operational risk.

• Partner with product and knowledge engineering teams to shape technical direction and the AI roadmap.

• Contribute to enterprise AI security practices, including prompt-injection protection, context-leak prevention, tenant isolation, access controls, and policy separation where applicable.

Requirements

• Professional AI/ML engineering experience with a demonstrated track record of shipping production LLM systems used by real users.

• Hands-on experience building and debugging multi-step, tool-calling agent workflows using LangGraph, LangChain, or an equivalent framework.

• Strong understanding of LLM evaluation, including representative datasets, regression testing, LLM-as-judge approaches, and/or human evaluation loops.

• Experience designing retrieval and context-assembly systems, with the ability to make informed decisions about what information to retrieve, how much to provide, and why.

• Demonstrated ownership of production systems through deployment, monitoring, troubleshooting, and incident response, including experience diagnosing and resolving failures or regressions.

• Strong judgment when working across multiple model providers and evaluating tradeoffs involving quality, latency, cost, context, and operational risk.

• Experience with Neo4j and Cypher or a comparable graph database is strongly preferred, along with the ability to quickly learn graph data modeling.

• Strong Python development skills and production experience with technologies such as FastAPI, asynchronous services, automated testing, observability, and maintainable software interfaces.

• Experience with graph schema evolution, embeddings, and operating live knowledge graphs is a plus.

• Familiarity with enterprise AI security, including prompt injection, context isolation, tenant separation, role-based access, and policy-layer controls is advantageous.

• Experience with Azure and hybrid search technologies such as Azure AI Search, Pinecone, MongoDB Atlas, pgvector, Elasticsearch, or similar platforms is a plus.

• Previous experience in B2B enterprise SaaS environments is strongly preferred.

• Ability to work effectively in a small, fast-moving team where ownership, adaptability, and independent technical decision-making are important.

• Must be based in the United States; relocation is not considered for this position.

• Visa sponsorship is not available for this role.

Benefits

• High-impact opportunity to solve complex AI reasoning problems in a safety-critical enterprise environment.

• Evaluation-first engineering culture focused on making AI quality measurable, observable, and continuously improvable.

• Direct exposure to customer needs and feedback from safety teams across energy, utilities, infrastructure, construction, and manufacturing.

• High level of technical ownership within a small, early-stage team.

• Close collaboration with technical leadership, product, and knowledge engineering teams.

• Opportunity to make consequential architectural and product decisions and see implementations reach customers quickly.

• Exposure to advanced LLM orchestration, knowledge graphs, retrieval systems, AI evaluation, and production AI operations.

• Remote role for candidates based in the United States.

How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

Work location

  • US

Related jobs

Ready to make a decision?

This role is either worth your time or it isn’t.

Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.

Nothing is submitted automatically. You choose what happens next.

About this listing

Published on Lever under the board identifier Jobgether, which is the name the employer’s own job board carries. RoleSprint has not verified the company’s registered or trading name, so it is shown exactly as published rather than tidied up.

RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.

Published 30 September 2026, last checked today. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.

More searches like this one

Browse all current openings

No credit card requiredStart free