Staff Modeling Architect

Neurophos - Sunnyvale, CA

Hiring: Staff Modeling Architect Company: Neurophos Location: Sunnyvale, CA Job Posted Time: 2026-09-10 13:57:16 Employment Type: Full-time / Hybrid Target Skills & Keywords : AI, C++, Event-Driven, FPGA, Hugging Face, LLM, Matplotlib, NumPy, ONNX, Pandas, PyTorch, Python, SOC, SystemVerilog, Transformers About the job Experience: •8+ years of experience in hardware modeling, functional modeling, performance modeling, performance simulation, or accelerator performance analysis used by architects, RTL, compiler and runtime, or silicon teams. Graduate research may count toward this. Required Skills: •We are seeking a staff-level modeling architect to build the path from a production model or application to two things: a performance and energy number Neurophos will stand behind, and a functional model that software can boot against before tape-out. •Bring up inference workloads as they ship, including dense and Mixture of Experts (MoE) transformers, attention and KV cache, expert routing, quantization, and hybrid/SSM models, plus retrieval, speech, vision, and recommendation workloads where they map onto the accelerator. •Bind Hugging Face and PyTorch workloads to the programming model and runtime, then run them on the functional model so that software and architecture are looking at the same behavior. •Co-design tiling, scheduling, the instruction set architecture (ISA), the SRAM and High Bandwidth Memory (HBM) hierarchy, network-on-chip (NoC) traffic, and multi-chip mapping across pipeline, tensor, and sequence parallelism, including collectives. •Run roofline and limiter analysis and design space exploration across microarchitecture options, resolving bottlenecks between the compiler view and the hardware. •Develop Python energy and latency models in NumPy, Pandas, and Matplotlib that cover operators, tiling, SRAM and HBM traffic, and optical GEMM and vector-unit time. •Implement bit-accurate C++ functional models of optical GEMM, SRAM vector processors, dataflow engines, and HBM, including narrow arithmetic, so software can begin bring-up before tape-out. •Contribute to the C++ event-driven simulation kernel itself, including coroutines, timed components, and traces, rather than only calling into it. Qualifications: •BS, MS, or PhD in Computer Engineering, Electrical Engineering, Computer Science, or equivalent practical experience. •Track record of shipping a model or study that another team depended on, whether architecture, compiler, customer, or silicon. •Judgment to pick the right method for a given question among roofline, limiter analysis, analytical performance models, trace-driven simulation, transaction-level modeling (TLM), and RTL simulation. •Strong grounding in computer architecture, microarchitecture, memory systems, and AI accelerators, whether GPU, TPU, NPU, or custom SoC. •Modern C++ (C++17 or later) for functional models, performance models, and simulation infrastructure. •Python for models, analysis, and plots, including NumPy, Pandas, and Matplotlib. •Demonstrated capacity to build an LLM or accelerator workload from a model card or paper, covering prefill and decode, MoE, GEMM tiling, and quantization. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!