Institute of Foundation Models

Inference Optimization Intern – Performance Modeling

Institute of Foundation Models · Sunnyvale, CA
Sunnyvale, CA Intern Closed
Applications are closed for this role. It was originally posted 2026-06-24. It’s no longer accepting applicants — see roles Institute of Foundation Models is still hiring for →, or browse the live openings below.
Type
Internship
Experience
0-1 yr

This intensive internship offers a unique opportunity to contribute to the development of a simulator and profiling framework for foundation model inference on NVidia GPUs.

Responsibilities include:

•
Develop analytical performance models for GPU kernels and inference workloads.

•
Build and validate a simulator to estimate theoretical hardware performance limits.

•
Compare measured kernel performance against architectural peak throughput.

•
Identify performance bottlenecks in compute, memory, communication, and scheduling.

•
Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute.

•
Investigate PTX and SASS code generation to understand low-level execution behavior.

•
Collaborate with researchers and engineers to optimize inference kernels for transformer-based models.

•
Evaluate utilization of Tensor Cores, memory bandwidth, caches, and instruction pipelines.

•
Design profiling methodologies for Hopper and Blackwell architectures.

•
Document findings and provide actionable recommendations for performance improvements.
Currently pursuing a degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, High-Performance Computing, or a related quantitative discipline.
•
Experience with CUDA programming and GPU kernel development.

•
Understanding of NVIDIA GPU architecture and memory hierarchy.

•
Familiarity with performance profiling tools such as Nsight Systems and Nsight Compute.

•
Knowledge of PTX, SASS, and low-level GPU execution.

•
Experience optimizing CUDA kernels for throughput and latency.

•
Understanding of roofline analysis, performance modeling, and hardware utilization metrics.

•
Experience with deep learning frameworks such as PyTorch or TensorFlow.

•
Strong programming skills in C++, CUDA, and Python.
•
Performance engineering mindset.

•
Strong analytical and debugging abilities.

•
Interest in AI systems, inference optimization, and hardware-software co-design.

•
Ability to work independently on research and engineering challenges.

•
Excellent written and verbal communication skills.

PyTorchTensorFlowPythonC++
O
Data & Eval Operations Program Manager
Sunnyvale, CA
Operations
$250K–$450K
D
Machine Learning Engineer — GPU Kernel
Sunnyvale, CA
Data & ML
$150K–$450K
D
Machine Learning Engineer — Pre-training
Sunnyvale, CA
Data & ML
$150K–$450K
See all live roles at Institute of Foundation Models →
C
Software Engineer, Inference AI/ML
CoreWeave Sunnyvale, CA
Engineering
$92K–$135K
E
Sr. Software Engineer – Generative AI & Assistants, ArcGIS Pro
Environmental Systems Research Institute Redlands, CA
Engineering
$123K–$202K
J
Campus AI Research Engineer – Research Automation
Jump Trading Chicago, IL
Engineering
$270K–$330K
O
Machine Learning Performance Engineer
Optiver Holding New York, NY
Engineering
$200K–$200K
See all Engineering roles →
Applications closed