Infrastructure engineer

Writer

$120K — $150K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years of experience in Infrastructure engineering, DevOps, or a similar role focused on high-availability production systems.
  • Deep expertise with AWS and containerization technologies like Docker and Kubernetes.
  • Proficiency in programming languages such as Python, Java, or Go for automation.
  • Knowledge of monitoring tools like Prometheus and Grafana.
  • Ability to challenge the status quo and identify systemic weaknesses in systems.
  • Strong communication and collaboration skills to connect with cross-functional teams.
  • A sense of ownership over mission-critical systems and a goal for peak performance.

Responsibilities

  • Build and automate AI-native approaches for infrastructure management using Python, Go, or similar languages.
  • Design scalable, fault-tolerant infrastructure solutions on AWS, GCP, or Azure.
  • Ensure the reliability and efficiency of WRITER's core services by maintaining stringent SLOs and Error Budgets.
  • Manage the observability stack for monitoring and logging of distributed systems.
  • Lead incident response and perform root cause analyses to improve system resilience.
  • Collaborate with product and engineering teams on system designs for reliability and scalability.

Benefits

  • Generous PTO, plus company holidays.
  • Comprehensive medical, dental, and vision coverage for employees and families.
  • Paid parental leave for all parents (16 weeks).
  • Support for fertility and family planning.
  • Early-detection cancer testing through Galleri.
  • Flexible spending account and dependent FSA options.
  • Health savings account with company contribution for eligible plans.
  • Annual wellness and learning development stipends.
  • Company-wide and team off-sites.
  • Competitive compensation, company stock options, and 401k.
Full Job Description
About the role

At WRITER, our mission to expand human capacity with super intelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7. As an Infrastructure engineer, you'll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows. This isn't just about keeping the lights on; it's about pushing the boundaries of what's possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI. You'll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need.

This is a hybrid position, based out of our New York City hub. You'll report to our director of engineering.

What you'll do
  • Use and build AI native approaches for operational tasks and infrastructure management and platforms using Python, Go, or similar languages, significantly reducing manual toil across our production environment
  • Design and implement scalable, fault-tolerant infrastructure AI solutions on public cloud providers (AWS, GCP, Azure) to support WRITER's rapidly expanding, high-traffic AI platform
  • Own the reliability, performance, and efficiency of WRITER's core services, defining and upholding stringent Service Level Objectives (SLOs) and Error Budgets
  • Own the observability stack for monitoring, logging, and alerting systems to ensure rapid detection of issues across our complex distributed systems
  • Lead incident response, post-mortems, and root cause analyses, applying learnings to proactively prevent future outages and build a more resilient system architecture
  • Collaborate closely with product and engineering teams, providing expert guidance on system design for reliability, performance, and scalability from conception through launch


☆ What you need
  • A solid 7+ years of experience in Infrastructure engineering, DevOps, Production engineering, Cloud platform or a similar role focused on building and operating large-scale, high-availability production systems
  • Deep expertise with cloud platforms (AWS strongly preferred), containerization technologies like Docker and Kubernetes, and Infrastructure-as-Code tools such as Terraform
  • Strong proficiency in programming languages such as Python, Java, Go for automation and monitoring
  • Knowledge of monitoring and logging tools (e.g., Prometheus, Grafana, ELK Stack) to maintain system health and performance
  • Demonstrated ability to Challenge the status quo, proactively identify systemic weaknesses, and propose innovative solutions to complex reliability problems
  • Excellent communication, collaboration, and problem-solving skills, with a talent for building strong relationships and Connecting with cross-functional teams
  • A strong sense of ownership and accountability, eager to Own mission-critical systems and drive them toward peak performance and unparalleled reliability


Benefits & perks (US Full-time employees)
  • Generous PTO, plus company holidays
  • Medical, dental, and vision coverage for you and your family
  • Paid parental leave for all parents (16 weeks)
  • Fertility and family planning support
  • Early-detection cancer testing through Galleri
  • Flexible spending account and dependent FSA options
  • Health savings account for eligible plans with company contribution
  • Annual work-life stipends for:
    • Wellness stipend for gym, massage/chiropractor, personal training, etc.
    • Learning and development stipend
  • Company-wide off-sites and team off-sites
  • Competitive compensation, company stock options and 401k

Similar Jobs

More Jobs at Writer

More Information Technology Jobs

Find similar Infrastructure engineer jobs: