Senior Platform Engineer
We are seeking a Senior Platform Engineer to build, administer, automate, secure, and operate enterprise Databricks and cloud data-platform environments. The role is responsible for Databricks workspace administration, Unity Catalog governance, identity and access management, compute and cluster policies, infrastructure automation, CI/CD enablement, monitoring, reliability, cost management, and production support.
The ideal candidate combines hands-on Databricks administration with strong AWS or Azure platform engineering, Terraform, Python, Kubernetes, CI/CD, security, networking, observability, and Site Reliability Engineering practices.
Key Responsibilities
Databricks Administration and Platform Engineering
• Administer Databricks accounts, workspaces, metastores, catalogs, schemas, external locations, storage credentials, connections, shares, recipients, and platform configurations.
• Provision and manage development, test, staging, and production workspaces using standardized, repeatable patterns.
• Configure workspace settings, repositories, jobs, notebooks, SQL warehouses, instance pools, job clusters, all-purpose compute, and serverless capabilities.
• Define and enforce cluster policies, approved runtime versions, libraries, init scripts, autoscaling, tagging, and compute-usage standards.
• Manage Databricks Runtime and platform upgrades, compatibility testing, release planning, maintenance windows, and rollback procedures.
• Support Databricks Workflows, Delta Live Tables or Lakeflow Declarative Pipelines, Databricks SQL, MLflow, model registry, feature engineering, and structured-streaming services.
• Troubleshoot workspace, permissions, connectivity, compute, storage, job execution, library, runtime, and performance issues.
• Maintain administration standards, platform documentation, knowledge articles, support procedures, and operational runbooks.
Unity Catalog, Identity, Security, and Governance
• Design and manage Unity Catalog metastores, catalogs, schemas, managed and external tables, volumes, external locations, storage credentials, and grants.
• Implement user and group provisioning through single sign-on, SCIM, identity-provider integration, and enterprise directory services.
• Administer account-level and workspace-level users, groups, service principals, permissions, entitlements, and access-control models.
• Implement role-based and attribute-based access controls, least-privilege permissions, separation of duties, and privileged-access procedures.
• Configure secure access patterns for secrets, tokens, credentials, service principals, private endpoints, storage, and external systems.
• Enable audit logging, lineage, system tables, tagging, data classification, row-level security, column masking, and compliance reporting.
• Partner with security and governance teams on encrypti on, key management, network controls, data loss prevention, retention, auditability, and regulatory requirements.
• Review access, monitor privileged activities, remediate policy violations, and support internal and external audits.
Cloud Infrastructure and Networking
• Build and operate Databricks on AWS or Azure, including secure integration with cloud storage, identity, networking, encryption, and monitoring services.
• On AWS, work with S3, IAM, KMS, VPC, PrivateLink, security groups, Route 53, CloudWatch, Secrets Manager, and related services.
• On Azure, work with ADLS Gen2, Microsoft Entra ID, managed identities, Key Vault, virtual networks, private endpoints, network security groups, Azure Monitor, and related services.
• Configure control-plane and data-plane connectivity, private networking, DNS, routing, firewall, proxy, and egress controls.
• Integrate Databricks with cloud data lakes, APIs, databases, message platforms, and enterprise applications.
• Support Kubernetes, Docker, EKS or AKS, and container-based platform services where required.
• Contribute to capacity planning, disaster recovery, high availability, backup, restoration, and business-continuity exercises.
Infrastructure as Code and Automation
• Build and maintain reusable Terraform modules using cloud and Databricks providers.
• Automate workspace, network, Unity Catalog, storage, identity, compute-policy, cluster, job, permission, and monitoring configurations.
• Manage Terraform state, workspaces, variables, modules, versioning, policy checks, drift detection, and controlled promotion across environments.
• Use Python, shell scripting, Databricks CLI, REST APIs, and SDKs to automate administrative and operational tasks.
• Implement self-service workspace, catalog, schema, access, and compute vending with appropriate approval and governance controls.
• Maintain configuration standards and reduce manual administration through repeatable automation.
DevOps, CI/CD, and Release Engineering
• Design and support CI/CD pipelines using GitHub Actions, GitLab CI/CD, Jenkins, Azure DevOps, or Harness.
• Automate deployment of notebooks, jobs, workflows, libraries, policies, infrastructure, and platform configuration.
• Support Databricks Asset Bundles, Git integration, artifact management, environment promotion, testing, approvals, and rollback.
• Integrate security scanning, policy validation, infrastructure testing, and release evidence into delivery pipelines.
• Enable engineering teams through templates, reusable pipelines, documentation, and self-service platform capabilities.
• Partner with application, data-engineering, and DevOps teams to ensure deploy ment standards are consistent and supportable.
Reliability, Monitoring, Operations, and FinOps
• Establish monitoring, alerting, dashboards, logs, metrics, traces, and health checks for Databricks and connected cloud services.
• Use CloudWatch, Azure Monitor, Datadog, Splunk, New Relic, or similar platforms to monitor availability, compute utilization, failures, security events, and cost.
• Define platform service-level indicators, service-level objectives, operational metrics, and error budgets.
• Lead incident response, problem management, root-cause analysis, corrective actions, and post-incident reviews.
• Manage vulnerability remediation, runtime patching, dependency updates, security exceptions, and platform lifecycle activities.
• Optimize cluster sizing, autoscaling, pools, SQL warehouses, job concurrency, serverless usage, storage, and workload scheduling.
• Implement budget controls, chargeback or showback tagging, utilization reporting, anomaly detection, and cost-optimization recommendations.
• Participate in operational support rotations and maintain escalation paths with Databricks and cloud providers.
• Coordinate platform upgrades, disaster-recovery tests, security reviews, and production-readiness assessments.
Collaboration and Technical Leadership
• Partner with architecture, data engineering, security, cloud, network, governance, FinOps, and service-management teams.
• Advise engineering teams on Databricks platform standards, secure patterns, deployment models, performance, and cost.
• Conduct technical reviews and ensure solutions meet enterprise architecture and operational-support requirements.
• Mentor platform engineers and administrators and lead knowledge-transfer sessions.
• Communicate platform health, risks, dependencies, incidents, and improvement roadmaps to technical and business stakeholders.
• Drive continuous improvement in automation, reliability, security, developer experience, and operational efficiency.
Required Qualifications
• Typically 710 years of cloud, DevOps, Site Reliability Engineering, infrastructure, or platform-engineering experience.
• At least 3 years of hands-on Databricks platform administration in an enterprise environment.
• Strong experience administering Databricks workspaces, Unity Catalog, compute, cluster policies, jobs, SQL warehouses, permissions, and service principals.
• Strong experience with AWS or Azure infrastructure, identity, storage, networking, encryption, monitoring, and private connectivity.
• Strong proficiency with Terraform and infrastructure-as-code practices.
• Experience with Python, shell scrip