SRE Monitoring Platform Software Engineer (Entry Level)

Bitdeer Technologies Group

$80K — $95K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 0-2 years of software engineering experience; fresh graduates with strong projects or internships are encouraged.
  • Solid fundamentals in a programming language (Go preferred), ability to write clean, tested, and readable code.
  • Understanding of computer science fundamentals including data structures, algorithms, and operating system concepts.
  • Basic knowledge of distributed systems principles such as caching and retries.
  • Exposure to monitoring tools like Prometheus or Grafana, and ability to write basic queries.
  • Familiarity with Linux systems and basic shell usage.
  • Kubernetes basics, including familiarity with Pods and Deployments.

Responsibilities

  • Build metrics and logs storage components, and code for data ingestion and queries.
  • Contribute to alert frameworks, including implementing and tuning default alert rules.
  • Help build services for cluster health and topology management.
  • Participate in the development of orchestration and scheduling components.
  • Instrument services for observability using OpenTelemetry, and create dashboards and runbooks.
  • Write unit and integration tests for delivered components, and engage in testing activities.

Benefits

  • Mentorship from senior engineers guiding your professional growth.
  • Exposure to a greenfield project with a well-defined architecture.
  • Opportunities to learn the full observability stack at production scale.
Full Job Description
Position Overview

Bitdeer is building an AI-operated GPU cloud - a global fleet of self-built and OEM-rented data centers running the world's most valuable compute, run by a platform that observes, protects, and operates the fleet. The SRE Platform team builds the monitoring and automation substrate that every other squad - storage, network, GPU, K8S, and L1 operators - depends on. Their signals become the system you help build.

As an entry-level Software Engineer on the SRE / Monitoring Platform team, you contribute to the NeoCloud SRE platform - the multi-region system that observes, protects, and operates a GPU rental fleet across self-built and OEM-rented data centers. You join a bounded context led by a senior engineer, take well-scoped components from design to production code that ships through GitOps + the CICD release pipeline, follows the Plugin Framework conventions, meets declared SLOs, and stays drift-free.

This is a build + learn role. You write code, write tests, and operate what you build under the guidance of a senior engineer. You participate in on-call as a shadow before taking primary. Within 12 months, you should be delivering components independently within your assigned area and growing toward owning a sub-context.

Key Responsibilities
    Where you'll contribute (guided by a senior engineer)
    • Collection + Storage - help build collection-agent, metrics-store / logs-store / traces-store / profiles-store, enrichment-service, collection-monitor. Write ingestion, query, and storage-path code.
    • Alert + Correlation + SLO - contribute to alert-engine-framework, alert-correlation, slo-framework; implement and tune default alert rules.
    • Topology + Cluster-Health - contribute to topology-service, cluster-health-rollup, OSS-SRE-tool collection plugins for K8s / Slurm / Ray / Volcano / Kueue / KubeRay.
    • Remediation + Workflow + Jobs - help build remediation-actuator, orchestration / workflow components, inspection probes, job-scheduler.
    • Observability instrumentation - instrument services with metrics, logs, and traces via OpenTelemetry; build dashboards; write runbooks an on-call can follow.
    • Test discipline - write unit / integration / contract tests for everything you ship; participate in chaos and soak tests led by senior engineers.

    Why this is a great first role
    • Greenfield with a well-defined vision. The Plugin Framework, GitOps pipeline, and SLO framework are decided; you build components inside them with a clear blueprint - not from a blank page.
    • You learn the full observability stack at production scale - ingest, query, storage - by building it, not just using it.
    • Mentorship-heavy. You work directly with senior and principal engineers who own the architecture; their expertise becomes your growth path.

Job Requirements
    • 0-2 years of software engineering experience (new graduates with strong projects or internships welcome).
    • Solid fundamentals in one programming language - Go (preferred), Python, Java, or Rust. You can write clean, tested, readable code and explain your design choices.
    • CS fundamentals - data structures, algorithms, concurrency, basic networking (TCP / HTTP), and operating-system concepts (processes, threads, I/O). You can reason about correctness and performance.
    • Distributed systems basics - you understand the ideas behind idempotency, retries, back-pressure, caching, and eventual consistency, even if you haven't operated them at scale yet. Eagerness to go deep.
    • Monitoring / observability exposure - some hands-on with Prometheus, Grafana, Loki, or similar; can write a basic PromQL query and instrument a service. Eagerness to learn the ingest, query, and storage path of a real observability stack.
    • Familiarity with Linux and the shell; comfort reading system logs and using standard debugging tools.
    • Kubernetes basics - understand Pods, Services, Deployments; have run something on K8s (a project, lab, or internship).
    • Git + CI basics - branching, pull requests, and have used a CI pipeline (GitHub Actions, GitLab CI, or similar).
    • Test discipline - you write unit and integration tests as a habit, not an afterthought.
    • Communication - clear written and verbal English; can write a good PR description and ask good questions.
    • Curiosity and a learning mindset - the most important qualifier. You're excited to learn GPU / AI infrastructure, AIOps, distributed systems, and observability at production scale.

Nice-to-Haves
    • Internship or project in monitoring / observability, telemetry pipelines, or platform / SRE tooling.
    • Exposure to GPU / AI-infra - DCGM, InfiniBand / RoCE, Kubernetes GPU Operator, Slurm / Ray. Interest counts more than depth.
    • Exposure to AIOps / ML-adjacent tooling (anomaly detection, alert correlation).
    • Contributions to open-source observability or cloud-native projects.

Similar Jobs

More Jobs at Bitdeer Technologies Group

More Information Technology Jobs

Find similar SRE Monitoring Platform Software Engineer (Entry Level) jobs: