Brak wyników spełniających kryteria wyszukiwania.
Founding Inference Engineer — FPGA + GPU Disaggregated Serving
alloycompute.aiGlobalnie
Founding Inference Engineer — FPGA + GPU Disaggregated Serving
AlloyCompute · Early-stage startup · Full-time · Significant equity
About Us
AlloyCompute.ai is building the next generation of LLM inference infrastructure on FPGAs. We run disaggregated inference pipelines that split workloads across FPGA and GPU — putting each phase of the serving stack on the silicon best suited for it — and we're already achieving sub-200ms end-to-end latency. We're early, funded, and moving fast. This is a founding-engineer role with meaningful equity and direct ownership of core architecture.
The Role
You'll own the design and optimization of our heterogeneous inference stack, where FPGA and GPU work as one system. You'll decide where prefill, decode, KV-cache management, and networking live, and you'll write the code that makes those boundaries disappear. You'll work directly with the founders and shape both the product and the technical culture.
What You'll Do
Architect and build disaggregated inference pipelines spanning FPGA and GPU, including prefill/decode splitting and KV-cache transfer across devices
Write and optimize custom CUDA kernels for attention, GEMM, quantization, and sampling paths
Extend and integrate open-source inference engines (vLLM, SGLang) with our FPGA backend — scheduling, batching, paged attention, speculative decoding
Design and implement FPGA dataflow in Verilog/HDL: systolic arrays, memory controllers, on-chip interconnect, and PCIe/Ethernet interfaces
Attack network and interconnect latency at every layer — RDMA/RoCE, NIC offload, kernel-bypass networking, and FPGA-to-GPU direct transfer
Profile, benchmark, and squeeze the stack: tokens/sec, time-to-first-token, tail latency, and cost per million tokens
What We're Looking For
Deep hands-on experience with CUDA kernel development and GPU performance engineering (Nsight, occupancy tuning, memory hierarchy optimization)
Real production experience with modern inference engines — vLLM, SGLang, TensorRT-LLM, or equivalent internals-level work
Strong FPGA skills: Verilog/SystemVerilog HDL, timing closure, HLS familiarity a plus, experience with AMD/Xilinx or Intel/Altera toolchains
Solid grasp of network latency engineering: RDMA, kernel bypass (DPDK/io_uring), NIC-level optimization, distributed serving topologies
Understanding of LLM serving internals: continuous batching, paged/radix attention, KV-cache management, quantization (FP8/INT4)
Comfort with ambiguity, bare-metal debugging, and shipping without a safety net
Nice to Have
Experience with disaggregated or heterogeneous serving architectures (e.g., Mooncake-style prefill/decode separation)
P4/SmartNIC or network-attached accelerator experience
Contributions to open-source inference projects
What We Offer
Founding-level equity — you're early enough for it to matter
Direct influence over architecture, roadmap, and hiring
Access to serious hardware: latest FPGAs, GPUs, and high-speed networking to experiment with
Competitive salary, flexible location, and a team that operates at the metal
Zainteresowany ofertą?
Aplikuj już teraz!