Powiązane oferty

Brak wyników spełniających kryteria wyszukiwania.

company logo

Founding Inference Engineer — FPGA + GPU Disaggregated Serving

alloycompute.aiIkona lokalizacjiGlobalnie

Żródlo publikacji: JustJoin.it
Rodzaj zatrudnienia
Rodzaj zatrudnieniaPełny etat
Doświadczenie
DoświadczenieSenior
Dodano
Dodano21 lipca 2026
Wykryte przez nas
Wykryte przez nas22 lipca 2026
Zarobki
Zarobki4377 - 13 130 EUR

Founding Inference Engineer — FPGA + GPU Disaggregated Serving


AlloyCompute · Early-stage startup · Full-time · Significant equity


About Us
AlloyCompute.ai is building the next generation of LLM inference infrastructure on FPGAs. We run disaggregated inference pipelines that split workloads across FPGA and GPU — putting each phase of the serving stack on the silicon best suited for it — and we're already achieving sub-200ms end-to-end latency. We're early, funded, and moving fast. This is a founding-engineer role with meaningful equity and direct ownership of core architecture.


The Role
You'll own the design and optimization of our heterogeneous inference stack, where FPGA and GPU work as one system. You'll decide where prefill, decode, KV-cache management, and networking live, and you'll write the code that makes those boundaries disappear. You'll work directly with the founders and shape both the product and the technical culture.
What You'll Do

  • Architect and build disaggregated inference pipelines spanning FPGA and GPU, including prefill/decode splitting and KV-cache transfer across devices

  • Write and optimize custom CUDA kernels for attention, GEMM, quantization, and sampling paths

  • Extend and integrate open-source inference engines (vLLM, SGLang) with our FPGA backend — scheduling, batching, paged attention, speculative decoding

  • Design and implement FPGA dataflow in Verilog/HDL: systolic arrays, memory controllers, on-chip interconnect, and PCIe/Ethernet interfaces

  • Attack network and interconnect latency at every layer — RDMA/RoCE, NIC offload, kernel-bypass networking, and FPGA-to-GPU direct transfer

  • Profile, benchmark, and squeeze the stack: tokens/sec, time-to-first-token, tail latency, and cost per million tokens

What We're Looking For

  • Deep hands-on experience with CUDA kernel development and GPU performance engineering (Nsight, occupancy tuning, memory hierarchy optimization)

  • Real production experience with modern inference engines — vLLM, SGLang, TensorRT-LLM, or equivalent internals-level work

  • Strong FPGA skills: Verilog/SystemVerilog HDL, timing closure, HLS familiarity a plus, experience with AMD/Xilinx or Intel/Altera toolchains

  • Solid grasp of network latency engineering: RDMA, kernel bypass (DPDK/io_uring), NIC-level optimization, distributed serving topologies

  • Understanding of LLM serving internals: continuous batching, paged/radix attention, KV-cache management, quantization (FP8/INT4)

  • Comfort with ambiguity, bare-metal debugging, and shipping without a safety net

Nice to Have

  • Experience with disaggregated or heterogeneous serving architectures (e.g., Mooncake-style prefill/decode separation)

  • P4/SmartNIC or network-attached accelerator experience

  • Contributions to open-source inference projects

What We Offer

  • Founding-level equity — you're early enough for it to matter

  • Direct influence over architecture, roadmap, and hiring

  • Access to serious hardware: latest FPGAs, GPUs, and high-speed networking to experiment with

  • Competitive salary, flexible location, and a team that operates at the metal

Zainteresowany ofertą?

Aplikuj już teraz!