AI Benchmark Engineer, Inference Performance

Silicon Data - United States

Hiring: AI Benchmark Engineer, Inference Performance Company: Silicon Data Location: United States Job Posted Time: 2026-09-17 04:57:11 Employment Type: Contract / Remote Target Skills & Keywords : Firmware, LLM, Linux, Node.js, Python, TensorRT, vLLM About the job Experience: •4+ years of engineering experience, with substantial hands-on work in LLM inference, model serving, or performance engineering. Required Skills: •This is a hands-on engineering role. You will write the harnesses, run the workloads, chase down the anomalies, and own the methodology that explains why the numbers are what they are. •Build and maintain the inference benchmarking harness. •Automated, containerized, reproducible runs across inference engines (vLLM, TensorRT-LLM, SGLang, TGI), model families, and hardware targets. •Own the inference metrics that matter. •Time-to-first-token, inter-token latency, sustained and peak throughput, latency percentiles under concurrency, goodput under SLO constraints, and their expression per dollar and per watt. •Design workload profiles that reflect real usage. •Input and output length distributions, concurrency patterns, streaming versus batch, prefill-heavy versus decode-heavy — rather than synthetic best-case runs. •Capture and validate the configuration fingerprint. Qualifications: •GPU profiling and telemetry experience — Nsight, DCGM, driver-level APIs, or equivalent. •Exposure to MLPerf Inference or comparable standardized benchmark suites. •Operational familiarity with non-NVIDIA accelerators (AMD, TPU, Cerebras, Groq, or custom silicon). •Background in certification, attestation, provenance, or audit-grade data products. •Understanding of GPU hardware economics, cloud compute pricing, or AI infrastructure procurement. •We are open to W-2 employment, onshore 1099 contract, or offshore corp-to-corp engagement — tell us which structure you are looking for and we will work with it. •You must be available during 8:00 AM – 1:00 PM Eastern Time, Monday through Friday. Outside that window, work when you work best. •We care about the quality and reproducibility of what you ship, not hours logged. Compensation: •Flexible work environment (work from home / hybrid options) •Competitive compensation and, for W-2 hires, meaningful equity Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!