AI Engineer, Inference

Firmus Technologies - Sydney, New South Wales, Australia

Hiring: AI Engineer, Inference Company: Firmus Technologies Location: Sydney, New South Wales, Australia Job Posted Time: 2026-09-17 06:42:17 Employment Type: Full-time Target Skills & Keywords : C++, CI/CD, GitOps, Hugging Face, Kubernetes, LLM, Load Balancing, Load Testing, Nim, Node.js, Python, RAG, TensorRT, Triton, vLLM About the job Experience: •5+ years of software engineering experience, including 3+ years in AI inference, model serving, ML systems, high-performance computing, distributed systems, or comparable performance-critical environments. •Demonstrated experience building, operating, or materially improving production model-serving platforms, inference APIs, GPU-backed services, AI developer platforms, or multi-tenant AI systems. •Applied hands-on capability in one or more modern inference frameworks, such as TensorRT-LLM, TensorRT, SGLang, vLLM, Triton Inference Server, NVIDIA Dynamo, NVIDIA NIM, Hugging Face Text Generation Inference, or equivalent technologies. •In-depth knowledge of the NVIDIA AI software stack, including CUDA, cuDNN, NCCL, TensorRT, GPU profiling, distributed communication, and GPU performance analysis. •Practical understanding of LLM and generative-AI serving behavior, including prompt processing, token generation, batching, context length, concurrency, KV-cache management, prefill and decode performance, request scheduling, model routing, and latency-throughput trade-offs. Required Skills: •Build, operate, and continuously improve self-hosted AI inference services for internal applications, customer-facing products, and future Inference-as-a-service offerings. •Define and implement standard model-onboarding workflows covering model intake, compatibility validation, packaging, runtime selection, optimization, deployment, endpoint registration, testing, release, and lifecycle management. •Provision and manage secure, scalable inference endpoints for common AI application patterns, including interactive generation, RAG, embeddings, reranking, batch processing, multimodal use cases, tool calling, and agentic workflows. •Develop reusable deployment templates, APIs, SDKs, configuration standards, and self-service workflows for users to request, configure, access, monitor, update, and retire model endpoints. •Interface directly with leading inference frameworks and toolkits, such as TensorRT-LLM, TensorRT, SGLang, vLLM, Triton Inference Server, NVIDIA Dynamo, NVIDIA NIM, CUDA, cuDNN, NCCL, and related serving, profiling, and observability tools. •Optimize model-serving performance using appropriate techniques, including quantization, compilation, batching, continuous batching, request routing, KV-cache management, prefix caching, speculative decoding, load balancing, model routing, memory optimization, and distributed parallelism. •Build and validate reusable inference recipes that specify compatible model versions, framework and runtime versions, precision formats, GPU configurations, topology requirements, scaling approaches, scheduler profiles, benchmark results, and expected performance envelopes. •Use quantization and optimization approaches such as NVFP4, FP8, INT8, TensorRT compilation, kernel optimization, efficient attention mechanisms, and memory-management techniques while maintaining agreed model-quality targets. Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!