Lead Site Reliability Engineer

Apply
Apply

Share

successfully icon

Successfully

The vacancy has been successfully added to favorites

location icon

Wilmington, US, Delaware, United States of America

specialization icon

Technical Support (SL3)

lob icon

BCM Industry

date icon

31/07/2026

Req. VR-124325

Apply
Project description

Responsible at the expert level for ensuring the reliability, scalability, performance, and operational excellence of critical banking platforms and applications. Serves as a senior individual contributor responsible for designing, implementing, and improving Site Reliability Engineering (SRE) practices across the software development lifecycle. Works closely with application development, infrastructure, platform engineering, and business teams to enhance system resiliency through automation, observability, testing, and proactive operational management while coaching and influencing others.

Responsibilities
bullet icon

Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure following enterprise technology standards and SRE best practices.

bullet icon

Lead initiatives to improve system reliability, availability, performance, and operational maturity through automation and engineering excellence.

bullet icon

Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical business services.

bullet icon

Develop comprehensive observability strategies leveraging Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting solutions.

bullet icon

Design and maintain end-to-end monitoring solutions that provide actionable insights into application, infrastructure, and customer experience health.

bullet icon

Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints.

bullet icon

Lead incident response activities for high-severity production events, coordinating cross-functional teams to restore services and minimize customer impact.

bullet icon

Perform and facilitate Root Cause Analysis (RCA) activities, ensuring corrective and preventive actions are identified, prioritized, and implemented.

bullet icon

Drive operational excellence through automation of repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls.

bullet icon

Partner with development teams to build reliable and observable services throughout the Software Development Lifecycle (SDLC).

bullet icon

Design, develop, and execute automated regression testing strategies to validate application stability, reliability, and performance following deployments and infrastructure changes.

bullet icon

Review test coverage and reliability validation approaches to ensure comprehensive testing and risk mitigation.

bullet icon

Create, maintain, and improve Infrastructure as Code (IaC) solutions using Terraform for cloud infrastructure provisioning, configuration management, and environment standardization.

bullet icon

Support and optimize Microsoft Azure environments, including Azure App Services, resource management, scaling strategies, deployment automation, and application lifecycle management.

bullet icon

Utilize Azure-native tools such as Azure Monitor, Application Insights, Log Analytics, and related services to improve platform visibility and reliability.

bullet icon

Drive implementation of performance testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness within assigned domains.

bullet icon

Establish operational readiness standards and ensure applications meet reliability, scalability, observability, and supportability requirements before production deployment.

bullet icon

Review architectural designs and provide recommendations to improve platform resiliency, operational efficiency, and cloud optimization.

bullet icon

Lead capacity planning, performance tuning, and workload optimization efforts across production environments.

bullet icon

Develop and maintain operational runbooks, incident playbooks, knowledge articles, and standard operating procedures.

bullet icon

Serve as a key partner with engineering, infrastructure, cybersecurity, architecture, and support teams to identify and implement continuous process improvements spanning organizational boundaries.

bullet icon

Communicate system health, reliability trends, operational risks, and remediation strategies to technical and business stakeholders.

bullet icon

Present reliability initiatives, operational metrics, and engineering recommendations at architecture reviews, technical forums, and leadership meetings.

bullet icon

Mentor engineers on observability, cloud engineering, automation, SRE principles, and operational best practices.

bullet icon

Understand and adhere to the Company's risk and regulatory standards, policies, and controls in accordance with the Company's Risk Appetite.

bullet icon

Identify reliability, operational, and technology risks requiring escalation to management.

bullet icon

Promote an environment that supports a culture of belonging and reflects the Client brand.

bullet icon

Maintain Client internal control standards, including timely implementation of internal and external audit findings and regulatory requirements as applicable.

bullet icon

Complete other related duties as assigned.

Skills

Must have

bullet icon

Strong experience in observability and monitoring, including hands-on expertise with:

bullet icon

Dynatrace

bullet icon

OpenTelemetry (OTel)

bullet icon

Distributed tracing

bullet icon

Metrics collection and analysis

bullet icon

Centralized logging and log aggregation

bullet icon

Alerting and dashboard development

bullet icon

Proven experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployments.

bullet icon

Strong proficiency in Infrastructure as Code (IaC) using Terraform.

bullet icon

Experience with CI/CD pipelines, deployment automation, and operational tooling.

bullet icon

Expert knowledge of production systems monitoring, incident management, and operational troubleshooting.

bullet icon

Strong understanding of application performance management, distributed systems, and modern cloud-native architectures.

bullet icon

Cloud & Platform Expertise

bullet icon

Strong experience with Microsoft Azure, including:

bullet icon

Azure App Services

bullet icon

Resource Groups

bullet icon

Azure networking concepts

bullet icon

Scaling and performance optimization

bullet icon

Deployment and release management

bullet icon

Application lifecycle management

bullet icon

Experience leveraging Azure-native operational tooling such as:

bullet icon

Azure Monitor

bullet icon

Application Insights

bullet icon

Log Analytics

bullet icon

Azure dashboards and alerting

bullet icon

Experience supporting cloud-native and hybrid infrastructure environments.

bullet icon

Reliability & Engineering Practices

bullet icon

Demonstrated experience implementing and operating SRE practices, including:

bullet icon

Service Level Objectives (SLOs)

bullet icon

Service Level Indicators (SLIs)

bullet icon

Error budgets

bullet icon

Incident management

bullet icon

Problem management

bullet icon

Root Cause Analysis (RCA)

bullet icon

Reliability automation

bullet icon

Ability to improve system reliability through:

bullet icon

Performance tuning

bullet icon

Capacity planning

bullet icon

Observability-driven insights

bullet icon

Proactive issue detection

bullet icon

Reliability engineering initiatives

bullet icon

Experience developing automated recovery mechanisms and self-healing solutions.

bullet icon

Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures.

Nice to have

bullet icon

Experience supporting large-scale enterprise applications in regulated environments.

bullet icon

Strong analytical and troubleshooting skills related to production systems and distributed architectures.

bullet icon

Experience working in Agile and DevOps operating models.

bullet icon

Ability to work autonomously and lead complex reliability initiatives.

bullet icon

Strong organizational and time management skills.

bullet icon

Advanced verbal and written communication skills.

bullet icon

Experience driving project milestones and delivery commitments.

bullet icon

Proven experience leading major incident response and post-incident improvement efforts.

bullet icon

Experience partnering with architecture, infrastructure, cybersecurity, and application development teams.

bullet icon

Experience with scripting and automation using PowerShell, Python, Bash, or similar technologies.

bullet icon

Industry certifications in Azure, Terraform, Cloud Engineering, or Site Reliability Engineering preferred.

Other
seniority icon

Languages

English: C1 Advanced

seniority icon

Seniority

Lead

Wilmington, US, Delaware, United States of America

Req. VR-124325

Technical Support (SL3)

BCM Industry

31/07/2026

Req. VR-124325

Apply for Lead Site Reliability Engineer in Wilmington, US, Delaware

*Indicates a required field

Under the terms of your specific consent or to perform our obligations under a contract with you, as applicable, we, Luxoft Holding Inc. will manually and electronically process your personal data, specifically your first name, last name, phone number, e-mail address and other data you provide us through this form.


Within this context, we process personal data only for the specific purpose(s) indicated in the individual consent language or other notices provided below.


We will – insofar as reasonably necessary for the purpose you have agreed to and within the scope of applicable laws – transfer your personal data to other entities within the Luxoft Group and to the group of third party recipients listed in our Privacy Notice. Such Recipients can be located outside the European Union (EU) and/or the European Economic Area (EEA) (“Third Countries”). The Third Countries concerned, e.g. the USA, may not have the level of data protection that you enjoy e.g. under the GDPR. This can result in disadvantages such as an impeded enforcement of data subjects’ rights, a lack of control over further processing and access by state authorities. You may only have limited legal remedies against this. Insofar our transfer of your personal data to recipients in Third Countries is not covered by an adequacy decision of the EU Commission, we achieve an adequate level of data protection as further detailed out in our Privacy Notice.


With your consent, we personalise marketing communications to you by way of carrying out marketing research analysis, analysing the surfing-behaviour of our website visitors and to adjust it to their detected tendencies, as well as to plan more efficient future marketing activities. This personalised marketing does not include any automated decision-making activities.


Further information on how we process personal data in general is available in our Privacy Notice. You may withdraw any given consent at any time. The withdrawal of your consent(s) will not affect the lawfulness of processing before its withdrawal. For any request in this context, please e-mail us at: DPO@luxoft.com.


Before uploading CV or any other information to this website, to learn more about your obligations and restrictions arising from the use of this website, please read our Terms of Use.