Position OverviewBitdeer is building an AI-operated GPU cloud - a global fleet of self-built and OEM-rented data centers running the world's most valuable compute, run by a platform that observes, protects, and operates the fleet. The SRE Platform team builds the monitoring and automation substrate that every other squad - storage, network, GPU, K8S, and L1 operators - depends on. Their signals become the system you help build.
As an entry-level Software Engineer on the SRE / Monitoring Platform team, you contribute to the NeoCloud SRE platform - the multi-region system that observes, protects, and operates a GPU rental fleet across self-built and OEM-rented data centers. You join a bounded context led by a senior engineer, take well-scoped components from design to production code that ships through GitOps + the CICD release pipeline, follows the Plugin Framework conventions, meets declared SLOs, and stays drift-free.
This is a build + learn role. You write code, write tests, and operate what you build under the guidance of a senior engineer. You participate in on-call as a shadow before taking primary. Within 12 months, you should be delivering components independently within your assigned area and growing toward owning a sub-context.
Key ResponsibilitiesWhere you'll contribute (guided by a senior engineer)
- Collection + Storage - help build collection-agent, metrics-store / logs-store / traces-store / profiles-store, enrichment-service, collection-monitor. Write ingestion, query, and storage-path code.
- Alert + Correlation + SLO - contribute to alert-engine-framework, alert-correlation, slo-framework; implement and tune default alert rules.
- Topology + Cluster-Health - contribute to topology-service, cluster-health-rollup, OSS-SRE-tool collection plugins for K8s / Slurm / Ray / Volcano / Kueue / KubeRay.
- Remediation + Workflow + Jobs - help build remediation-actuator, orchestration / workflow components, inspection probes, job-scheduler.
- Observability instrumentation - instrument services with metrics, logs, and traces via OpenTelemetry; build dashboards; write runbooks an on-call can follow.
- Test discipline - write unit / integration / contract tests for everything you ship; participate in chaos and soak tests led by senior engineers.
Why this is a great first role
- Greenfield with a well-defined vision. The Plugin Framework, GitOps pipeline, and SLO framework are decided; you build components inside them with a clear blueprint - not from a blank page.
- You learn the full observability stack at production scale - ingest, query, storage - by building it, not just using it.
- Mentorship-heavy. You work directly with senior and principal engineers who own the architecture; their expertise becomes your growth path.
Job Requirements- 0-2 years of software engineering experience (new graduates with strong projects or internships welcome).
- Solid fundamentals in one programming language - Go (preferred), Python, Java, or Rust. You can write clean, tested, readable code and explain your design choices.
- CS fundamentals - data structures, algorithms, concurrency, basic networking (TCP / HTTP), and operating-system concepts (processes, threads, I/O). You can reason about correctness and performance.
- Distributed systems basics - you understand the ideas behind idempotency, retries, back-pressure, caching, and eventual consistency, even if you haven't operated them at scale yet. Eagerness to go deep.
- Monitoring / observability exposure - some hands-on with Prometheus, Grafana, Loki, or similar; can write a basic PromQL query and instrument a service. Eagerness to learn the ingest, query, and storage path of a real observability stack.
- Familiarity with Linux and the shell; comfort reading system logs and using standard debugging tools.
- Kubernetes basics - understand Pods, Services, Deployments; have run something on K8s (a project, lab, or internship).
- Git + CI basics - branching, pull requests, and have used a CI pipeline (GitHub Actions, GitLab CI, or similar).
- Test discipline - you write unit and integration tests as a habit, not an afterthought.
- Communication - clear written and verbal English; can write a good PR description and ask good questions.
- Curiosity and a learning mindset - the most important qualifier. You're excited to learn GPU / AI infrastructure, AIOps, distributed systems, and observability at production scale.
Nice-to-Haves- Internship or project in monitoring / observability, telemetry pipelines, or platform / SRE tooling.
- Exposure to GPU / AI-infra - DCGM, InfiniBand / RoCE, Kubernetes GPU Operator, Slurm / Ray. Interest counts more than depth.
- Exposure to AIOps / ML-adjacent tooling (anomaly detection, alert correlation).
- Contributions to open-source observability or cloud-native projects.