Opportunity Overview
We are looking for a hands-on Principal Azure Platform & Cloud Operations Architect to assess
and improve how we run critical production systems on Azure. You will evaluate our current
Cloud Operations processes and platform architecture, identify automation and improvement
opportunities, implement stronger operational patterns, and act as the escalation SME when the
team hits technical roadblocks—especially across AKS, networking, and deployments.
Key Responsibilities
workflows (provisioning, deployments, incident response, change management) and
implement process corrections, automation opportunities, and operational guardrails to
reduce manual effort and improve reliability.
maturity of Azure Kubernetes Service (AKS) environments, including cluster topology,
node pools, upgrades, scaling, resiliency patterns, ingress/egress, workload identity,
secrets, and runtime security.
in Azure networking and connectivity patterns (VNET design, routing/UDRs, NSGs,
DNS, private endpoints, firewalls, load balancers, gateways, and secure egress/ingress)
and troubleshoot complex network and performance issues impacting production
systems.
Terraform with reusable modules, clear lifecycle management, environment consistency,
and safe change practices.
using Argo CD and Helm, improving repeatability, release confidence, environment
promotion, and rollback strategies.
Jenkins / GitHub Actions) to implement quality gates, validation, security scanning, and
automated delivery patterns to production.
alerting, and dashboards using Azure Monitor, Log Analytics, and Application Insights to
create actionable signals and reduce noise; promote production readiness practices
(runbooks, readiness reviews, operational checklists).
point for high-severity incidents, guiding triage and recovery, leading root cause analysis,
and ensuring corrective/preventive actions are implemented through automation and
platform improvements.
hands-on guidance during complex technical challenges, raising overall capability and
establishing consistent engineering standards.
teams to align operational patterns, platform guardrails, and production readiness across
services and environments.
Qualifications
Bachelor's or master's degree in Computer Science, Engineering, or a related field.
8+ years of hands-on experience in cloud platform engineering, DevOps/SRE, or cloud
operations, with ownership of production-grade systems.
and deep experience running Kubernetes in production (upgrades, scaling, failure
modes, troubleshooting).
diagnose complex multi-layer issues across AKS + network + application boundaries.
environment consistency, safe rollout practices).
and a strong understanding of release strategies and operational controls.
Actions) and improving deployment reliability through automated checks and gates.
Application Insights) to improve service health visibility and incident response
effectiveness.
operational tooling and reduce repetitive manual work.
pressure and guide teams throug
Salary: $80 – $90 / hour