SRE Monitoring Platform Software Engineer (Entry Level)

Bitdeer Technologies Group

$90K — $110K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 0-2 years of software engineering experience, including strong projects or internships.
  • Solid coding skills in one programming language (Go preferred, also Python, Java, or Rust).
  • Understanding of CS fundamentals: data structures, algorithms, concurrency, and networking basics.
  • Familiarity with distributed systems concepts such as idempotency and caching.
  • Experience with monitoring/observability tools like Prometheus or Grafana, including basic query writing.
  • Proficient with Linux and shell commands, including system log analysis.
  • Basic understanding of Kubernetes and experience using CI/CD pipelines.

Responsibilities

  • Contribute to building components for metrics storage, logs, traces, and monitoring.
  • Implement and tune alert rules in collaboration with senior engineers.
  • Develop services for topology management and cluster health monitoring.
  • Assist in creating remediation tools and job scheduling workflows.
  • Instrument applications for observability using metrics, logs, and traces via OpenTelemetry.
  • Write comprehensive tests for all code and participate in chaos testing processes.

Benefits

  • Greenfield project with a clear, well-defined vision and goals.
  • Opportunity to learn the complete observability stack through hands-on experience.
  • Mentorship from senior and principal engineers who guide your professional growth.
Full Job Description
Position Overview

Bitdeer is building an AI-operated GPU cloud - a global fleet of self-built and OEM-rented data centers running the world's most valuable compute, run by a platform that observes, protects, and operates the fleet. The SRE Platform team builds the monitoring and automation substrate that every other squad - storage, network, GPU, K8S, and L1 operators - depends on. Their signals become the system you help build.

As an entry-level Software Engineer on the SRE / Monitoring Platform team, you contribute to the NeoCloud SRE platform - the multi-region system that observes, protects, and operates a GPU rental fleet across self-built and OEM-rented data centers. You join a bounded context led by a senior engineer, take well-scoped components from design to production code that ships through GitOps + the CICD release pipeline, follows the Plugin Framework conventions, meets declared SLOs, and stays drift-free.

This is a build + learn role. You write code, write tests, and operate what you build under the guidance of a senior engineer. You participate in on-call as a shadow before taking primary. Within 12 months, you should be delivering components independently within your assigned area and growing toward owning a sub-context.

Key Responsibilities
    Where you'll contribute (guided by a senior engineer)
    • Collection + Storage - help build collection-agent, metrics-store / logs-store / traces-store / profiles-store, enrichment-service, collection-monitor. Write ingestion, query, and storage-path code.
    • Alert + Correlation + SLO - contribute to alert-engine-framework, alert-correlation, slo-framework; implement and tune default alert rules.
    • Topology + Cluster-Health - contribute to topology-service, cluster-health-rollup, OSS-SRE-tool collection plugins for K8s / Slurm / Ray / Volcano / Kueue / KubeRay.
    • Remediation + Workflow + Jobs - help build remediation-actuator, orchestration / workflow components, inspection probes, job-scheduler.
    • Observability instrumentation - instrument services with metrics, logs, and traces via OpenTelemetry; build dashboards; write runbooks an on-call can follow.
    • Test discipline - write unit / integration / contract tests for everything you ship; participate in chaos and soak tests led by senior engineers.

    Why this is a great first role
    • Greenfield with a well-defined vision. The Plugin Framework, GitOps pipeline, and SLO framework are decided; you build components inside them with a clear blueprint - not from a blank page.
    • You learn the full observability stack at production scale - ingest, query, storage - by building it, not just using it.
    • Mentorship-heavy. You work directly with senior and principal engineers who own the architecture; their expertise becomes your growth path.

Job Requirements
    • 0-2 years of software engineering experience (new graduates with strong projects or internships welcome).
    • Solid fundamentals in one programming language - Go (preferred), Python, Java, or Rust. You can write clean, tested, readable code and explain your design choices.
    • CS fundamentals - data structures, algorithms, concurrency, basic networking (TCP / HTTP), and operating-system concepts (processes, threads, I/O). You can reason about correctness and performance.
    • Distributed systems basics - you understand the ideas behind idempotency, retries, back-pressure, caching, and eventual consistency, even if you haven't operated them at scale yet. Eagerness to go deep.
    • Monitoring / observability exposure - some hands-on with Prometheus, Grafana, Loki, or similar; can write a basic PromQL query and instrument a service. Eagerness to learn the ingest, query, and storage path of a real observability stack.
    • Familiarity with Linux and the shell; comfort reading system logs and using standard debugging tools.
    • Kubernetes basics - understand Pods, Services, Deployments; have run something on K8s (a project, lab, or internship).
    • Git + CI basics - branching, pull requests, and have used a CI pipeline (GitHub Actions, GitLab CI, or similar).
    • Test discipline - you write unit and integration tests as a habit, not an afterthought.
    • Communication - clear written and verbal English; can write a good PR description and ask good questions.
    • Curiosity and a learning mindset - the most important qualifier. You're excited to learn GPU / AI infrastructure, AIOps, distributed systems, and observability at production scale.

Nice-to-Haves
    • Internship or project in monitoring / observability, telemetry pipelines, or platform / SRE tooling.
    • Exposure to GPU / AI-infra - DCGM, InfiniBand / RoCE, Kubernetes GPU Operator, Slurm / Ray. Interest counts more than depth.
    • Exposure to AIOps / ML-adjacent tooling (anomaly detection, alert correlation).
    • Contributions to open-source observability or cloud-native projects.

Similar Jobs

More Jobs at Bitdeer Technologies Group

More Information Technology Jobs

Find similar SRE Monitoring Platform Software Engineer (Entry Level) jobs: