Senior Software Engineer, Infrastructure

Stream

$155K — $200K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in infrastructure, platform, DevOps, or SRE engineering with a focus on infrastructure
  • Strong coding skills in Go or Python, with a proven software engineering background
  • Experience with Kubernetes at meaningful production scale, including cluster strategy and workload design
  • Demonstrated ability to optimize cloud costs or efficiency on AWS or GCP
  • Expertise in managing high-scale, high-load production systems
  • Strong understanding of cloud fundamentals, including networking and IAM
  • Comfortable in a small team environment, capable of multifaceted contributions

Responsibilities

  • Design and operate infrastructure for real-time systems with millions of connections
  • Direct Kubernetes cluster design and workload migration processes
  • Re-architect workloads for cost and performance during the AWS to GCP migration
  • Lead initiatives for cloud cost management and efficiency improvements
  • Develop production-quality Go and Python code for internal tools and automation
  • Conduct post-migration performance tuning and capacity planning
  • Collaborate with engineers across teams on service design and reliability

Benefits

  • Work with an exceptional team of engineers
  • Engagement in open-source software projects
  • Generous PTO starting at 20 days annually
  • Company equity as part of compensation
  • Fitness stipend offered
  • Provided with a Macbook Pro for work
  • Budget allocated for Learning and Development
  • Opportunities to attend global conferences and meetups
  • Possibility of visiting offices in Boulder, CO, and Amsterdam, NL
Full Job Description
Senior Software Engineer, Infrastructure

The role

We are hiring a Senior Software Engineer to help rebuild the platform underneath Stream. Over the next year the infrastructure team is moving from AWS to GCP, moving onto Kubernetes, and relocating 35 to 40 Postgres shards off managed RDS to self-hosted, while the platform keeps serving billions of API requests a month. You will own parts of that outright.

This is a small, senior team without the support structures of a large organisation. You will write code most of the time and make infrastructure calls on your own. Success looks like systems that scale predictably under load, cloud spend that falls per unit of traffic, and migrations that land without incident.

This is a full-time job opening based in Toronto (3 days hybrid).

What you will do
  • Design, build and operate infrastructure for real-time systems carrying millions of concurrent connections and billions of monthly API requests.
  • Drive Kubernetes end to end: cluster architecture, workload design and the migration of existing services. You will be designing clusters, not operating someone else's.
  • Re-architect workloads as part of the AWS to GCP migration, for cost and performance rather than a lift and shift.
  • Own cloud cost and efficiency work: find the levers, measure them against real spend and utilisation data, and show what moved.
  • Write production Go and Python: internal services, platform tooling and automation that change how product and SDK engineers deploy, observe and debug.
  • Lead post-migration tuning and capacity planning, closing the loop between the architecture you chose and what production actually does.
  • Work with backend, video and moderation engineers on system design, reliability targets and tradeoffs that cross service boundaries.
  • Take part in on-call, incident response and root cause analysis, and turn what you find into durable fixes.
What we are looking for
  • 5+ years in infrastructure, platform, DevOps or SRE engineering, with clear depth in infrastructure over application development.
  • A software engineering background. You have built systems, not only configured them. Production coding experience in Go or Python. Scripting-only backgrounds are not a fit.
  • Kubernetes at meaningful production scale, past operations: you have driven cluster strategy, designed workloads, or led a migration, and you have tuned what came out the other side for cost and efficiency.
  • Cloud cost or efficiency optimisation you personally led on AWS or GCP, with an outcome you can put a number on. FinOps practice is a plus.
  • Direct experience running high-scale, high-load production systems.
  • Strong cloud fundamentals across networking, compute, storage and IAM, and the habit of asking why a system behaves the way it does instead of accepting the default.
  • Comfortable in a small team: leading a project and reviewing a PR in the same week.
  • AI tooling already in your engineering workflow. Applied use, not familiarity.
Bonus points
  • Both AWS and GCP, and migration experience between providers.
  • PostgreSQL at scale: sharding, replication strategy, partitioning tradeoffs, ideally self-hosted.
  • Real-time systems: WebSockets, WebRTC, streaming or other persistent-connection workloads.
  • The wider stack: CockroachDB, Redis, Terraform, and a Prometheus-based observability stack.
  • An API-first or infrastructure company at scaleup stage.
  • Open source contributions to infrastructure or platform tooling.
  • Writing or talks on cloud, platform or distributed systems.
  • Formal FinOps practice, or owning cloud commitment and reservation strategy.
  • Work on developer-facing API or SDK products.
Our stack
  • Go, gRPC, RocksDB, Python
  • PostgreSQL, RabbitMQ
  • GCP
  • Grafana, Prometheus, ELK (Elasticsearch and Kibana)
  • Jaeger and Tempo for distributed tracing, Datadog
  • Redis, Memcached
  • Claude Code, Cursor
You will thrive here if
  • You want infrastructure problems at a scale most engineers never touch, and the autonomy to own them.
  • You ship fast and learn fast, including when it is hectic.
  • You are self-directed and comfortable working with a globally distributed team across time zones.
You probably will not if
  • You want tightly scoped tickets and step-by-step direction.
  • You need a calm, highly predictable environment.
  • You would rather wait for a defined process than act.
Compensation and benefits

Stream employees enjoy some of the best job benefits in the industry:
  • A team of exceptional engineers
  • The chance to work on OSS projects
  • 20 days of PTO first year of employment, 24 days starting your second year
  • Company equity
  • Fitness stipend
  • A Macbook Pro provided
  • A Learning and Development budget
  • The opportunity to attend or present to global conferences and meetups
  • The possibility to visit our offices in Boulder, CO and Amsterdam, NL

Salary Range: CA$155,000 to CA$200,000 per year, plus stock options. Final offer within this range depends on experience and interview outcome.

Hybrid office policy: applicants based (or relocating to) one of our office locations are expected to work according to the applicable local office attendance policy.

Note for external recruiters: We currently have this role covered and do not accept unsolicited agency resumes. We are not responsible for any fees related to unsolicited resumes.

Similar Jobs

More Jobs at Stream

More Information Technology Jobs

Find similar Senior Software Engineer, Infrastructure jobs: