🌍 Global Opportunities
⚑ Updated Hourly
πŸŽ“ Student Friendly
⏰

parttimejobs.work

Flexible Work, Better Balance

⏰ Full-time

AI Inference Performance Engineer

NVIDIA
Location πŸ“ Santa Clara, United States
Posted πŸ“… June 03, 2026
Work Type ⏰ Full-time

Position Overview

We optimize and benchmark GenAI inference on NVIDIA's latest accelerators, defining the industry’s performance standards across language models, video generation, and speech workloads. We work directly within TensorRT-LLM, SGLang, and vLLM, building the tools that evaluate serving performance at scale. This team sits at the intersection of GPU performance engineering and public accountability.


What You Will Be Doing:
+ Drive industry benchmark results: own the end-to-end optimization pipeline, implement and integrate optimizations in quantization, scheduling, memory management, and distributed inference across TensorRT-LLM, SGLang, and vLLM.
+ Define and optimize cutting-edge workloads: identify and shape next-generation inference benchmarks, multi-turn coding, agentic workflows, and other emerging AI use cases. Collaborate with framework and kernel teams to push performance to its extreme on large-scale LLM-MoE models, vision-language models, video diffusion mode...

Apply Now

Submit Application β†’

Quick and easy application process

Job Details

⏰
Employment Type
Full-time
πŸ“Š
Category
other-general
🏠
Work Arrangement
On-site
πŸ“
Location
Santa Clara, United States