Sr. Software Engineer - Platform Performance & Resilience (AI-Enabled)

Toshiba Global Commerce Solutions - External

$100K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 4-6+ years of software engineering experience.
  • Strong proficiency in Node.js and Java.
  • Experience in performance engineering or distributed systems architecture.
  • Design experience with determining timeouts and concurrency controls.
  • Solid understanding of SLOs and non-functional validation.
  • Experience deploying services in Kubernetes environments.
  • Skilled in debugging and profiling distributed systems.

Responsibilities

  • Design mechanisms for transaction integrity across various systems.
  • Define failure-mode strategies for connectivity issues and data conflicts.
  • Engineer patterns to prevent cascading system failures.
  • Establish latency budgets and performance metrics.
  • Collaborate to eliminate performance bottlenecks pre-production.
  • Develop AI-enabled automated resilience validation systems.
  • Architect telemetry systems for improved transaction traceability.

Benefits

  • Group health coverage (medical, dental, & vision)
  • Employee Assistance Programs
  • Pre-tax spending accounts
  • 401(k) plan with company match
  • Life insurance
  • Pet Insurance
  • Employee discounts
  • Generous paid holiday schedule and vacation days
Full Job Description
Toshiba Global Commerce Solutions is seeking a Senior Software Engineer - Platform Performance & Resilience that plays a key role in engineering performance, resilience, and observability across a three-tier distributed architecture spanning edge devices, in-store servers, and cloud services. This role uses AI-enabled automation to validate and enforce production-grade reliability, with the ultimate goal of delivering measurable system stability at retail scale. The position operates at the intersection of distributed systems architecture, performance engineering, reliability validation, and intelligent automation.

This role is a hybrid role in Durham, NC office. On-site interviews will be required.

Responsibilities:

Architect Reliability Across Edge-Store-Cloud
  • Design and implement platform mechanisms that ensure transaction integrity and availability across POS terminals, store middleware, and cloud services.
  • Define and validate failure-mode strategies for intermittent connectivity, tier isolation, data replay, and synchronization conflicts.
  • Engineer patterns that prevent cascading failures and support graceful degradation under real-world load.

Engineer Performance at Retail Scale
  • Define latency budgets and performance envelopes across all tiers.
  • Build systems that measure and validate throughput, concurrency limits, and resource saturation.
  • Collaborate with development teams to eliminate bottlenecks before production.

Build Automated Resilience Validation
  • Develop AI-enabled systems that automatically generate and execute performance and resilience validation scenarios.
  • Integrate non-functional quality gates into CI/CD workflows.
  • Continuously evaluate timeout, retry, circuit breaker, and backoff strategies under stress.

Elevate Observability & Signal Quality
  • Architect structured telemetry across edge, store, and cloud tiers.
  • Ensure end-to-end transaction traceability.
  • Improve root-cause detection by strengthening monitoring signal-to-noise ratio.

Own Engineering Outcomes End-to-End
  • Produce technical designs and failure-mode analyses.
  • Implement and deploy platform components in Node.js and companion services in Java.
  • Drive production-readiness improvements based on performance data.

Required Qualifications:
  • 4-6+ years of professional software engineering experience.
  • Strong proficiency in Node.js and Java.
  • Proven experience in performance engineering, reliability engineering, or distributed systems architecture.
  • Demonstrated experience designing systems with deterministic timeouts, retry/backoff strategies, circuit breakers, and concurrency controls.
  • Experience modeling multi-tier systems (edge, middleware, cloud).
  • Solid understanding of SLOs, SLIs, and non-functional validation.
  • Experience deploying services in Kubernetes-based cloud environments.
  • Strong debugging and profiling skills for distributed systems.

Preferred Qualifications:
  • Experience building automated resilience or fault-injection systems.
  • Familiarity with event-driven architectures (Kafka, Pub/Sub, MQ).
  • Experience implementing structured observability frameworks.
  • Exposure to AI-enabled automation or workflow orchestration.
  • Experience optimizing systems in intermittently connected environments.


Toshiba Global Commerce Solutions, Inc. offers a competitive salary and generous benefits package including the following:

  • Group health coverage (medical, dental, & vision)
  • Employee Assistance Programs
  • Pre-tax spending accounts
  • 401(k) plan (with company match)
  • Company provided life insurance
  • Pet Insurance
  • Employee discounts
  • Generous paid holiday schedule, paid vacation & sick/personal days


Similar Jobs

More Jobs at Toshiba Global Commerce Solutions - External

More Information Technology Jobs

Find similar Sr. Software Engineer - Platform Performance & Resilience (AI-Enabled) jobs: