Skip to content

Jobgether

AI Research Engineer (Kernel & Inference Optimization)

Location
Spain, ES
Arrangement
Remote
Employment type
Full-time
Posted
30 September 2026 (today)

Checked todayApplications go to the employer, never to RoleSprint

About this role

Accountabilities • Design and deploy advanced model-serving architectures optimized for high throughput, low latency, and efficient memory utilization.

• Develop inference pipelines capable of operating effectively across diverse environments, including resource-constrained mobile devices and edge platforms.

• Establish clear performance targets covering response latency, token generation speed, throughput, memory footprint, and reliability.

• Build and execute controlled inference benchmarks in simulated and production environments, tracking latency, throughput, memory consumption, and error rates.

• Create and maintain representative datasets and simulation scenarios for evaluating model performance under real-world and resource-constrained conditions.

• Identify computational and memory bottlenecks across inference pipelines and implement solutions involving batching, networking, memory management, and other system-level optimizations.

• Develop custom GPU kernels and compute shaders for mobile hardware, including solutions written in Metal Shading Language (MSL).

• Apply advanced inference optimization techniques such as pruning, quantization, Flash Attention, KV caching, and speculative decoding.

• Design and optimize distributed inference systems using approaches such as tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU workloads.

• Work with cross-functional engineering and research teams to integrate optimized inference frameworks into production and edge-device applications.

• Define evaluation methodologies, document experimental results, compare performance against established benchmarks, and continuously refine optimization strategies.

• Monitor production performance and use empirical research to identify opportunities for further improvements in scalability, efficiency, and reliability.

Requirements:

• Degree in Computer Science or a related technical field; a PhD in NLP, Machine Learning, or a related discipline is highly relevant, particularly with a strong AI research track record and publications at leading conferences.

• Proven expertise in Metal Shading Language (MSL), including the ability to write custom compute shaders from scratch.

• Demonstrated experience with low-level kernel optimization and inference optimization on mobile or other resource-constrained devices.

• Track record of delivering measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications.

• Deep understanding of modern model-serving architectures, inference engines, and optimization techniques for high-performance AI deployment.

• Strong experience writing GPU kernels for mobile devices such as smartphones.

• Practical experience developing and deploying end-to-end inference pipelines, from model optimization through production integration on constrained hardware.

• Strong ability to apply empirical research and systematic experimentation to solve latency, computational, and memory challenges.

• Experience designing robust evaluation and benchmarking frameworks for inference systems.

• Knowledge of distributed inference techniques, including tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU clusters.

• Deep understanding of the mathematical foundations and architecture of diffusion models and Vision Transformers.

• Familiarity with modern inference optimization techniques including pruning, quantization, Flash Attention, KV Cache optimization, and speculative decoding such as EAGLE.

• Strong analytical and problem-solving abilities, with an ability to investigate complex system bottlenecks and turn research findings into practical engineering solutions.

• Excellent English communication skills and the ability to collaborate effectively with distributed, cross-functional technical teams.

Benefits:

• Opportunity to work on advanced AI systems spanning model serving, inference optimization, mobile computing, edge deployment, and large-scale distributed inference.

• Remote-first working environment with an international team.

• Exposure to cutting-edge AI research and practical systems engineering challenges.

• Opportunity to contribute to performance-critical infrastructure where improvements can have a measurable impact on real-world AI applications.

• Collaborative environment combining research-driven experimentation with hands-on engineering.

• Opportunity to work with advanced model architectures including diffusion models, Vision Transformers, and multimodal systems.

• Access to challenging technical problems involving GPU kernels, inference engines, memory optimization, and distributed computing.

How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

Work location

  • ES

Related jobs

Ready to make a decision?

This role is either worth your time or it isn’t.

Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.

Nothing is submitted automatically. You choose what happens next.

About this listing

Published on Lever under the board identifier Jobgether, which is the name the employer’s own job board carries. RoleSprint has not verified the company’s registered or trading name, so it is shown exactly as published rather than tidied up.

RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.

Published 30 September 2026, last checked today. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.

Browse all current openings

No credit card requiredStart free