Back to jobs
DataOps Consultant, Lakehouse Support
Successfully
Req. VR-123002
We are hiring a Support Engineer who combines production support expertise with DevOps engineering capability and AI fluency. This role spans the full D&E portfolio: in-house .NET/C# applications, APIs, event-driven systems, data platforms, and BI tools. You will ensure operational stability while continuously improving reliability through automation, observability, and modern DevOps practices.
You will work directly alongside Software Engineers as the bridge between development and operations, owning incident resolution, CI/CD pipeline health, infrastructure automation, and production readiness. You will leverage AI tools daily to accelerate triage, automate runbooks, and proactively detect issues before they impact users.
We are an AI-first engineering team: you will use AI-powered observability, automated diagnostics, and intelligent incident management to deliver faster resolution and higher system reliability.
Perform deep operational diagnostics across data pipelines (batch/stream);
Coordinate resolution across platform engineering, Data Engineers and data domain teams (Databricks, Snowflake, Palantir);
Ensure correct usage of ITSM tools (e.g., ServiceNow) for incident tracking and lifecycle management;
Monitor critical workflows using Dynatrace;
Workflows orchestration using StoneBranch;
Leverage metadata and documentation for troubleshooting and root cause analysis;
Maintain clear ownership, escalation paths and communication during incidents;
Reduce incident noise and recurring issues through structured problem management
Must have
Production Support and Incident Management:
Own L2/L3 production support across in-house applications (.NET/C#, Python), APIs, event-driven systems, and data platforms.
Lead incident triage, diagnosis, and resolution with structured root cause analysis (RCA) and post-incident reviews.
Manage incident, problem, and change management processes aligned with ITIL best practices.
Coordinate with development teams, infrastructure, and vendors for escalation and resolution of complex production issues.
Maintain and improve production verification suites, canary checks, and reliability playbooks.
DevOps and CI/CD:
Own and maintain CI/CD pipelines in Azure DevOps: build, test, deploy, and release automation for both application and data platform projects.
Infrastructure as Code: Terraform, Bicep, or ARM templates for provisioning and managing Azure cloud resources.
Container orchestration: Docker and Kubernetes (AKS) for deploying, scaling, and troubleshooting containerized services.
Environment management: provision, configure, and maintain development, staging, and production environments.
Implement and maintain deployment strategies (blue-green, canary, rolling) to minimize production risk.
Cloud and Infrastructure:
Azure: Practical experience with App Services, AKS, Functions, Event Hubs, Storage, Key Vault, Azure SQL, and Azure Monitor.
Manage cloud resource health, cost optimization, and security posture.
Support microservice architectures: service discovery, load balancing, scaling, and resilience patterns.
Networking fundamentals: DNS, load balancers, firewalls, VNets, and private endpoints in Azure.
Observability and Monitoring:
Implement and manage monitoring stacks: Dynatrace, Azure Monitor or equivalent.
Build and maintain alerting rules, dashboards, and SLO/SLI tracking for critical services and data pipelines.
Distributed tracing with OpenTelemetry: correlate trace IDs across services, diagnose latency bottlenecks, and validate signal fidelity.
Data observability: monitor data pipeline health, detect anomalies, latency issues, and data loss using platforms like Monte Carlo Data or equivalent.
Log aggregation and analysis for rapid incident diagnosis.
Application and Data Platform Support:
Applications: Support .NET/C# backend services, REST/GraphQL APIs, event-driven systems (Solace, Kafka, RabbitMQ), and WPF/web front-ends.
Data Platforms: Support ETL/ELT pipelines (PySpark, Spark), big data platforms (Databricks, Palantir Foundry), and data warehouses.
Databases: Advanced SQL for diagnosing data issues, validating transformations, and supporting migrations (SQL Server, MongoDB).
Scheduling and Orchestration: Manage batch workloads in StoneBranch, Airflow, or equivalent.
AI-First Engineering (How We Work)
AI-powered incident triage: use AI tools to cluster alerts, suggest root causes, and recommend resolution steps from historical incident data.
Automated runbooks: build and maintain AI-assisted diagnostic scripts that accelerate common troubleshooting workflows.
AIOps: leverage AI/ML for anomaly detection, predictive alerting, and capacity forecasting across applications and data pipelines.
AI coding assistants (GitHub Copilot, Cursor) used daily for scripting, automation, configuration, and documentation.
Support AI-based applications and services deployed by D&E teams; understand model serving, inference pipelines, and their operational requirements.
Responsible AI: follow engineering playbooks for data handling (no sensitive data in prompts), security, and observability of AI-generated changes.
Scripting and Automation:
Python scripting for operational automation, data validation, and diagnostic tooling.
PowerShell for Windows administration and Azure resource management.
Build self-healing mechanisms and automated remediation where feasible.
Nice to have
Palantir (Foundry) — DQC setup in workshop, diagnostics of build failures and data pipelines
Monte Carlo — data quality alert management and observability
Azure DevOps — repository management, CI/CD
Working knowledge of investment management data: portfolio terminology (maturity date, value date), trade lifecycle, market and risk data; familiarity with domain systems (Optima, Cardia)
Understanding of data lineage, observability and metadata-driven operations concepts
Exposure to enterprise-scale data platforms and batch orchestration / data workflow tools
Languages
English: C1 Advanced
Seniority
Senior
Abu Dhabi, United Arab Emirates
Req. VR-123002
Technical Support (SL2)
BCM Industry
31/07/2026
Req. VR-123002
Apply for DataOps Consultant, Lakehouse Support in Abu Dhabi
*Indicates a required field