Tidy3D is a GPU-accelerated electromagnetic simulation product delivered as a cloud service. Behind the solver sits the platform that makes everything work: the Python client, the API, the job submission and scheduling layer, and the web application.
We are hiring a Software Engineer for the Tidy3D infrastructure team to own that platform. You will build new capability, keep the existing system healthy, and run the release and deployment process.
ResponsibilitiesDesign and build backend services for the Tidy3D platform.- Develop and operate the control plane: the APIs, services, and data model behind task submission, job state, and result delivery.
- Build scheduling and resource management for simulation jobs across heterogeneous GPU capacity.
- Handle the operational concerns that come with a multi-tenant product: authentication, authorization, usage metering, and quota enforcement.
Help with the deployment of our products to customers.- Manage packaging and release to PyPI, and keep client and backend versions compatible across a long tail of installed versions.
- Own the release pipeline end to end: versioning, CI/CD, staged rollout, and rollback.
- Standardize deployment patterns so the same product ships to our cloud, to customer-managed cloud accounts, and to on-premises installations.
Keep production healthy.- Instrument the platform and own its monitoring and alerting.
- Respond to incidents and debug across boundaries, from a customer's Python traceback down to a stuck job on a GPU node.
- Manage cloud cost and capacity as usage grows.
Requirements- 2+ years building and operating production cloud services. New grads with exceptional background are encouraged to apply too.
- Strong Python. You have written backend services in it, not just scripts.
- Hands-on experience with a major cloud provider such as AWS, Docker, and Kubernetes.
- Infrastructure as code in production, ideally Terraform, with reusable modules rather than hand-managed environments.
- Linux fluency and comfort operating in production.
Nice to have
- GPU or HPC workloads, cluster schedulers such as Slurm, Ray, or Kueue, or high-throughput batch compute.
- Designing, packaging, and distributing a developer-facing Python library or SDK.
- On-premises, self-hosted, or air-gapped software deployment, and the enterprise requirements that come with it: SSO, network isolation, security review.
- A background in physics, engineering, or scientific computing, or prior work on technical software for technical users.
- Open source contributions to infrastructure or scientific computing projects.
Benefits- Competitive compensation with equity of a fast-growing startup.
- Medical, dental, and vision health insurance.
- 401(k) Contribution.
- Gym allowance.
- Friendly, thoughtful, and intelligent coworkers.