Embedding VC

Member of Technical Staff - Efficient ML

Embedding VC · San Francisco, CA
San Francisco, CA Closed
Applications are closed for this role. It was originally posted 2026-01-15. It’s no longer accepting applicants — see roles Embedding VC is still hiring for →, or browse the live openings below.
Type
Full-time
Experience
8+ yr

Introducing Moonlake, AI for creating world simulations.

SCOPE OF WORK

Training efficiency

  • Dataloaders, fusion, activation remat, gradient checkpointing.
  • FSDP/ZeRO/tensor+pipeline parallel; NCCL tuning.

GPU + kernel performance

  • Nsight profiling, Triton/CUDA kernels, fused ops.
  • Flash-attention–style speedups, sequence packing, KV-cache tricks.

Inference optimization

  • Low-latency serving, continuous batching, speculative decoding.
  • Quantization (GPTQ/AWQ), distillation, pruning.

Infra + reliability

  • SLURM/K8s multi-node jobs, checkpoint hygiene.
  • Determinism, env pinning, GPU failure handling.

We are committed to being an on-site, in-person team currently based in San Mateo

E
Growth Engineer - Globalization
San Francisco, CA
Engineering
$300K–$400K
O
Member of Technical Staff, Infra
Mountain View, CA Hybrid
Operations
$140K–$200K
O
Member of Technical Staff, Platform
Mountain View, CA Hybrid
Operations
$120K–$200K
See all live roles at Embedding VC →
L
Member of Technical Staff - ML Scientist, Japanese Multimodal
Liquid AI Tokyo, Japan Hybrid
Operations
B
Member of Technical Staff - Atlas
Basis New York, NY
Operations
$100K–$300K
C
Agent Runtime & Systems - Member of Technical Staff
Callosum Technologies London, UK
Operations
$101K–$192K
C
Member of Technical Staff - RL Environments
Cohere London, UK Hybrid
Operations
$295K–$535K
See all Operations roles →
Applications closed