FuriosaAI

Senior Software Engineer - Inference Engine

FuriosaAI · Seoul, South Korea
Seoul, South Korea Closed
Applications are closed for this role. It was originally posted 2026-08-10. It’s no longer accepting applicants — see roles FuriosaAI is still hiring for →, or browse the live openings below.
Type
Full-time
Experience
3+ yr

ABOUT THE JOB

Software Engineer (Inference Engine) is responsible for developing and optimizing a high-performance inference engine for Large Language Models (LLMs) and multimodal LLMs running on FuriosaAI NPUs.

In this role, you will proactively research and apply the state-of-the-art inference optimization techniques to our inference engine. You will work in close collaboration with the compiler and hardware teams to enhance the engine's performance to its full potential.

RESPONSIBILITIES

  • Design and implement FuriosaAI’s next-generation inference engine for large and multimodal language models—comparable in capability to frameworks such as vLLM and SGLang—optimized for throughput, latency, and memory efficiency.
  • Design and implement advanced inference optimizations—such as speculative decoding, KV-cache management, tensor/model parallelism, memory-efficient execution, and scheduling—in our production inference engine.
  • Design and develop capabilities for distributed and scalable inference, including prefill–decode (PD) and encode–prefill–decode (EPD) disaggregation, disaggregated speculative decoding, and hierarchical and external KV-cache storage such as HiCache and Mooncake.
  • Collaborate closely with the Compiler team to co-design and optimize execution for FuriosaAI NPUs, improving system-level throughput, latency, and memory utilization.
  • Proactively research, evaluate, and integrate state-of-the-art inference optimization techniques and key features of LLM serving frameworks into our production inference engine.

MINIMUM QUALIFICATIONS

  • BS degree in Computer Science, Engineering, or a related field, with at least 3 years of relevant industry experience, or equivalent practical experience
  • Proficiency in Rust or C++ programming skill
  • Knowledge and passion of deep learning, LLM, and/or generative AI models
  • Excellent problem-solving and data analysis skills.
  • Strong communication and collaboration skills.

PREFERRED QUALIFICATIONS

  • Experience in building inference serving systems for large models, encompassing batching, scheduling, caching, and load balancing.
  • A deep understanding of performance optimization systems.
  • Proficiency in C++/CUDA or Triton kernel development
  • Contributions to open-source inference frameworks such as vLLM, SGLang, or TensorRT-LLM.

CONTACT

LLMRustC++
A
Senior Software Engineer - Maritime Integrated Solutions
Anduril Industries Costa Mesa, California, United States
Engineering
$191K–$253K
C
Senior Software Engineer - Network Platforms
Cloudflare Hybrid Hybrid
Engineering
$68K–$91K
C
Senior Software Engineer I, Inference
CoreWeave Sunnyvale, CA
Engineering
$139K–$204K
A
Senior Research Engineer
AssemblyAI New York, NY Remote
Engineering
$270K–$310K
See all Engineering roles →
Applications closed