Embedding VC

Member of Technical Staff - ML Infrastructure & Performance

Embedding VC · San Mateo, CA
San Mateo, CA Closed
Applications are closed for this role. It was originally posted 2025-12-12. It’s no longer accepting applicants — see roles Embedding VC is still hiring for →, or browse the live openings below.
Type
Full-time
Experience
8+ yr

Introducing Moonlake, AI for creating real-time interactive content

Mission: Improve Throughput, Latency, & Cost - deploying our models 2–10× faster & cheaper without quality regressions.

Scope of Work:

  • GPU performance: CUDA/Triton kernels, FlashAttention family, paged attention, CUDA Graphs.
  • Serving stack: TensorRT-LLM/Triton Inference Server, vLLM/TGI; continuous batching; on-GPU KV reuse; speculative decoding/medusa; mixture-of-agents routing.
  • Parallelism: FSDP/ZeRO, TP/PP/expert parallel; NCCL tuning.
  • Quantization/PEFT: AWQ/GPTQ/FP8; LoRA/DoRA serving.
  • Systems: Ray/k8s/Argo, observability (Prom/Grafana/OpenTelemetry), autoscaling, A/B infra, canary + rollback.

Tech signals:

Previous experience at Infra-heavy startups such as Databricks, Roblox

We are committed to being an on-site, in-person team currently based in San Mateo

LLMDatabricks
E
Growth Engineer - Globalization
San Francisco, CA
Engineering
$300K–$400K
E
Senior/Staff Frontend Engineer
San Francisco, CA Remote
Engineering
$100K–$200K
E
Senior/Staff Product Engineer
San Francisco, CA Remote
Engineering
$100K–$200K
See all live roles at Embedding VC →
N
Senior Member Technical Staff (MTS 3) - Machine learning
Nielsen Holdings Bangalore, India
Operations
P
Member of Technical Staff
Perplexity AI Belgrade
Operations
X
Member of Technical Staff - RL Inference
xAI Palo Alto, CA
Operations
$180K–$440K
L
Member of Technical Staff - Applied ML, Japanese Multimodal
Liquid AI Tokyo, Japan Hybrid
Operations
See all Operations roles →
Applications closed