Institute of Foundation Models

Inference Optimization Intern – Performance Modeling

Institute of Foundation Models · Sunnyvale, CA
Sunnyvale, CA Intern Posted 2026-06-24
Type
Internship
Experience
0-1 yr

This intensive internship offers a unique opportunity to contribute to the development of a simulator and profiling framework for foundation model inference on NVidia GPUs.

Responsibilities include:


Develop analytical performance models for GPU kernels and inference workloads.


Build and validate a simulator to estimate theoretical hardware performance limits.


Compare measured kernel performance against architectural peak throughput.


Identify performance bottlenecks in compute, memory, communication, and scheduling.


Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute.


Investigate PTX and SASS code generation to understand low-level execution behavior.


Collaborate with researchers and engineers to optimize inference kernels for transformer-based models.


Evaluate utilization of Tensor Cores, memory bandwidth, caches, and instruction pipelines.


Design profiling methodologies for Hopper and Blackwell architectures.


Document findings and provide actionable recommendations for performance improvements.
Currently pursuing a degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, High-Performance Computing, or a related quantitative discipline.

Experience with CUDA programming and GPU kernel development.


Understanding of NVIDIA GPU architecture and memory hierarchy.


Familiarity with performance profiling tools such as Nsight Systems and Nsight Compute.


Knowledge of PTX, SASS, and low-level GPU execution.


Experience optimizing CUDA kernels for throughput and latency.


Understanding of roofline analysis, performance modeling, and hardware utilization metrics.


Experience with deep learning frameworks such as PyTorch or TensorFlow.


Strong programming skills in C++, CUDA, and Python.

Performance engineering mindset.


Strong analytical and debugging abilities.


Interest in AI systems, inference optimization, and hardware-software co-design.


Ability to work independently on research and engineering challenges.


Excellent written and verbal communication skills.

PyTorchTensorFlowPythonC++
E
Eval360 - Error Analysis Engineer
Sunnyvale, CA
Engineering
$150K–$450K
D
Research Scientist - Vision Language Model
Sunnyvale, CA
Data & ML
$150K–$450K
D
Research Scientist, Agentic Data & Benchmarking
Sunnyvale, CA
Data & ML
$150K–$450K
See all 10+ roles at Institute of Foundation Models →
C
Staff Python / PyTorch Developer — Frontend Inference Compiler – Dubai
Cerebras Systems Europe Remote
Engineering
C
Software Engineer, Inference AI/ML
CoreWeave Sunnyvale, CA
Engineering
$92K–$135K
D
Principal System Software Engineer, AI Inference Execution
d-Matrix Santa Clara, CA Hybrid
Engineering
$195K–$285K
E
Sr. Software Engineer – Generative AI & Assistants, ArcGIS Pro
Environmental Systems Research Institute Redlands, CA
Engineering
$123K–$202K
See all Engineering roles →

Interested in this role?

Apply directly on the company site — no recruiter middleman, no account required.

Apply now →
Apply on company site