The application window is expected to close on: 10/29/2026
Job posting may be removed earlier if the position is filled or if a sufficient number of applications are received.
Meet the Team
We run the platform that serves foundational models to Cisco IT. Our Foundational Model Service, gives engineering teams across the company access to small, large language, and embedding models. youWe serve those models on Kubernetes clusters, with Nim, Vllm and other runtimes. Beyond serving, We benchmark, evaluate, monitor, and release new models as improvements and demand warrant. Our customers depend on the platform under a 99.9% uptime SLA, and we build and operate accordingly.
As a Software Engineer on the FMS team you'll keep that platform running and make it easier to run. You'll support model onboarding and releases, take your turn on call, help automate the validation that gates every deployment, and contribute to the monitoring that shows us and our customers how the service is behaving. You'll also apply Agentic solutions to our own operations to catch problems earlier and remediate them automatically.
Your Impact
You'll develop software consistent with Cisco Design Thinking Principles, with simplification and user experience at its core, using secure coding practices, protecting user privacy, and following software development best practices. You'll partner with design, product management, and other engineering teams to build the right solution for our customers. You'll create technical design documentation for the team, contribute to the documentation end users rely on, and debug platform issues both in development and in production.
This is a strong role for a systems or infrastructure engineer who has worked with LLMs, whether through academic coursework or by building and running them on the job, and who understands how inference serving works. We're looking for a self starter who takes on unfamiliar work and grows into it. You'll work with engineers who have built these systems from the ground up. New runtimes and new models arrive frequently, so the platform is always evolving, and you'll develop a deep understanding of developing AI infrastructure, model runtimes, and delivering products at scale. On call is shared across the team.
Minimum Qualifications:
• Bachelor's degree in Computer Science, Information Systems, or a related field with 3 years of related experience; master's degree with 1 year of related experience; PhD; or equivalent practical work experience.
• Experience developing and operating production services on Kubernetes and Linux, including exposure to GPU based AI infrastructure.
• Hands on experience serving models with vLLM, NVIDIA NIM, Triton, or a comparable inference runtime, including building supporting services.
• Experience leading technical work, mentoring engineers, documenting systems, and supporting critical services through an on call rotation.
Preferred Qualifications:
• Experience evaluating and benchmarking models to support production release decisions.
• Experience operating GPU workloads with NVIDIA CUDA or AMD ROCm.
• Familiarity with distributed inference, disaggregated serving, KV cache aware routing, capacity planning, or cost optimization.
• Experience fine tuning transformer models or evaluating RAG and agent systems using relevant frameworks.