Senior Software Engineer

Upbound

$150K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in software engineering, specifically in cloud services and Kubernetes.
  • Strong debugging skills in distributed systems with observability tools such as Prometheus and Grafana.
  • Experience building controllers that interact with the Kubernetes API server.
  • Proven ability to communicate effectively with customers to resolve complex technical issues.
  • Demonstrated ownership and problem-solving skills in production environments.

Responsibilities

  • Build and operate Upbound Spaces in production and troubleshoot multi-tenant SaaS environments.
  • Develop high-demand features that enhance user experience based on customer feedback.
  • Investigate and debug complex customer issues, focusing on performance bottlenecks and resource reconciliation.
  • Author detailed design documents and post-incident reviews to facilitate system improvements.
  • Manage the full project lifecycle for scalable cloud services including design and operational support.
  • Write and maintain Go code for Kubernetes API interactions, emphasizing observability and operational excellence.
  • Deploy and troubleshoot Kubernetes services in production environments using monitoring tools.

Benefits

  • Opportunity to make a meaningful impact in product and operational engineering.
  • Work with a growing team focused on technology innovation in the cloud.
  • Chance to contribute to open-source projects such as Crossplane.
  • Flexible remote working arrangements available for candidates in designated US markets.
Full Job Description
Upbound is hiring a Senior Software Engineer to help us build and operate Upbound Spaces, the multi-control-plane management software at the heart of the Upbound Platform. As part of the Spaces team, you'll help us scale Upbound to reliably support thousands of control planes while extending enterprise-grade control plane management and operations across both cloud and on-premises environments. Our team is growing, and this is an opportunity to make a meaningful engineering impact across both product development and production operations. For this particular opening, we're focused on growing our engineering presence in the San Francisco Bay Area and Seattle, and are prioritizing candidates who are currently based in, or able to work from, one of those two markets. **What You'll Do** - Actively build and operate Upbound Spaces in production, troubleshooting and resolving issues across multi-tenant SaaS environments, as well as contributing to Upbound's open-source projects, including Crossplane. - Take ownership of building features in high demand by Upbound's customers and deliver new functionality that will delight and amaze our users. - Investigate and debug complex issues in customer environments, including multi-control plane scenarios, resource reconciliation problems, and performance bottlenecks. - Communicate through thoughtful and thorough design documents for new initiatives and detailed post-incident reviews that drive system improvements. - Support the full project lifecycle for highly scalable and reliable services running in a cloud environment - discovery, analysis, architecture, design, review, documentation, building, migration, automation, deployment, production-readiness, and ongoing operational support. - Write and maintain Go code that interfaces with the Kubernetes API, such as operators, controllers, add-ons, etc., with a focus on observability, debuggability, and operational excellence. - Deploy, manage, and troubleshoot our Kubernetes services in production, using metrics, logs, and traces to identify and resolve issues quickly. - Build and maintain operational tooling for debugging customer environments, analyzing control plane health, and automating incident response. - Author documentation, user guides, runbooks, and blog posts to support and promote new features that you release. - Support the software release cycle for Spaces self-hosted distributions, including diagnosing issues in customer-managed deployments. - Participate in on-call rotation to support Upbound Cloud, responding to incidents and driving them to resolution. **What You'll Bring** - Have experience operating production cloud services at scale: monitoring, alerting, incident response, post-mortems, and continuous improvement of service reliability. - Have strong debugging skills across distributed systems, including experience with observability tools (Prometheus, Grafana, OpenTelemetry, distributed tracing) and techniques for diagnosing issues in production environments. - Have experience building and operating controllers that interact with the Kubernetes API server, including troubleshooting reconciliation loops, managing API rate limits, and optimizing controller performance. - Are comfortable working directly with customers to understand, reproduce, and resolve complex technical issues in their environments. - Take responsibility and ownership for solving problems even if they are outside your lane, especially during incidents affecting customer workloads. - Demonstrate excellence in your work, constantly trying to improve your skills and the operational posture of the systems you build. - Have empathy for customers and keep them in mind as you build solutions, understanding that reliability and debuggability are features. - Realize the importance of clear communication and effective collaboration to work as a team, deliver great results, and support customers through technical challenges. - Help create a safe environment where everyone can contribute, learn from failures, share on-call knowledge, and help each other grow as operators and engineers. #LI-REMOTE

Similar Jobs

More Jobs at Upbound

More Information Technology Jobs

Find similar Senior Software Engineer jobs: