Dale-Kurt Murray
hello@dalekurtmurray.com · New York, NY · linkedin.com/in/dalekurt · github.com/dalekurt
Summary
Senior Site Reliability Engineer and technical lead with 20+ years in infrastructure, specializing in incident management, disaster recovery, and large-scale cloud platforms. Experienced in leading platform incidents on production on-call, fixing failures at the root, and cutting about $1.5 million a year in database spend. Core expertise in Kubernetes, AWS, Terraform, Prometheus, and Grafana. Builds AI agent systems outside work, including an incident platform where an AI agent triages alerts and pages a human with a draft incident report.
Work Experience
Adobe – New York, NY
Senior Software Engineer / Senior Site Reliability Engineer (February 2026 – Present)
Platform and reliability engineering for Frame.io across 33 AWS accounts, 43 Kubernetes clusters, and 6 CockroachDB Cloud clusters. On the production on-call rotation, responsible for leading platform incidents.
- Scaling Team
- Optimized CockroachDB Cloud clusters, saving about $1.5 million annually, and designed a sizing strategy that cuts vCPU and storage costs without giving up horizontal or vertical scaling.
- Launched and co-developed a tool that captures production traffic using Istio EnvoyFilters and replays it in lower environments. The team now uses it to validate CockroachDB migration behavior, and its pod-scoped egress blocking isolates load tests on production data.
- Created the Crossplane composite resource and Terraform (S3 backups, VPC endpoints, Route 53 private DNS) behind preview environments, with capacity-aware placement across a cluster pool.
- Integrated preview clusters with Teleport via AWS PrivateLink using TLS certificates from Vault, replacing VPNs and direct credentials, and fixed five bugs in the Teleport auto-approval flow so engineers who aren’t on call get point-in-time restore access without waiting on SRE.
- Traced a preview database pool’s restore failures to bugs in two separate codebases and delivered five fixes in under a week, so the pool now heals itself.
Software Engineer / Site Reliability Engineer (November 2021 – February 2026)
- Search Team
- Crafted per-pull-request preview environments for OpenSearch, reducing setup time by 80%.
- Spearheaded the team’s Elastic Cloud to Amazon OpenSearch migration without any data loss.
- Automated backup and recovery for business continuity, and wrote the runbook the team follows.
- Implemented Claude on Amazon Bedrock and resolved IAM policy conflicts with AWS support.
- Platform Team
- Automated the AMI lifecycle with Packer, Python, and GitHub Actions (EKS, ECS, EMR, and GPU templates, Goss tests, CIS hardening, approval gates), and rotated AMI releases in EKS with Karpenter.
- Wrote the runbooks for rotating TLS certificates and IAM keys for third-party applications that can’t use IRSA, which other engineers follow and improve during on-call.
Frame.io (acquired by Adobe in 2021) – New York, NY
Senior Site Reliability Engineer (April 2020 – November 2021)
- Designed and built real-time malware scanning for uploaded media, moving ClamAV from EC2 to AWS Lambda with prewarming, concurrency tuning, automated signature updates, and Ghostscript protection.
- Enforced Aqua Security scanning policies with the security team, and added Snyk to CI/CD.
- Planned and executed upgrades of Aurora PostgreSQL and EKS in all environments with minimal downtime.
- Designed S3 and RDS disaster recovery, copying snapshots to a separate account against ransomware.
- Evaluated gVisor, Firecracker, and Kata Containers for isolated, short-lived execution environments.
DevOps Engineer (March 2018 – April 2020)
First member of the Union Square DevOps team, owning release strategy, CI/CD, and secrets management.
- Served as the technical lead for four DevOps engineers, on-site and remote, from Q3 2018 to Q2 2019.
- Ran Kubernetes on AWS and Google Cloud (GKE) with Helm and zero-downtime blue/green deployments.
- Built a highly available HashiCorp Vault cluster and added Twistlock and Snyk scanning to CI/CD.
Earlier experience. Cloud architect and Linux systems administrator for New York startups such as AdTheorent, iNet Giant, Keeps America, and Lead Liaison. Web operations manager at Simply Intense Media in Trinidad and Tobago, and DevOps engineer at Bakari Digital Labs in Kingston, Jamaica. Career began in Jamaica at Flow Jamaica (Columbus Communications), Restaurants of Jamaica, and Nutrition Products.
Projects
- Agent control plane: This project turns tickets into merged pull requests across nine repositories, including its own. It runs coding agents as transient Kubernetes jobs on Hatchet, with concurrency and budget controls, automated checks, and human approval before merging.
- Incident platform: One platform for outages and security incidents. An AI agent triages alerts, enriches them, and pages a human with a draft incident report, acting under its own RBAC role through typed, audited, reversible actions, with human approval for production changes. Built on Temporal and Kubernetes.
- Homelab: Talos Kubernetes cluster on Proxmox, linked with a Hetzner burst node over WireGuard. Argo CD keeps dev and production environments in sync from Git and creates preview environments for pull requests. Terraform provisions the nodes, Cloudflare, and GitHub, while OPA policies validate manifests during CI, with local LLMs running on a dedicated GPU node.
Skills
- Cloud: AWS (EC2, S3, IAM, EKS, ECS, EMR, Lambda, RDS Aurora, OpenSearch, Bedrock), Google Cloud (GKE, Compute Engine, load balancing, storage), Proxmox VE
- Networking: VPC, PrivateLink, VPC endpoints, Route 53 private DNS, Istio, Envoy, Kong Gateway, TLS
- Kubernetes: Helm, Crossplane, Argo CD, Karpenter, Talos Linux, Docker
- Infrastructure as code and CI/CD: Terraform, OpenTofu, Packer, GitHub Actions, CircleCI, Git
- Data and workflows: CockroachDB, Aurora PostgreSQL, Temporal, Hatchet
- Observability: Grafana, Prometheus, Datadog
- Security: HashiCorp Vault, Teleport, Aqua Security, Snyk, OPA
- AI: Amazon Bedrock, Claude Code, Ollama
- Languages: Python, Go, Bash
Certifications
- AWS Certified AI Practitioner, Amazon Web Services (January 2026, valid through January 2029)
- Certified Kubernetes Administrator (CKA), CNCF (January 2020, expired January 2023)
- AWS Certified Solutions Architect – Associate, Amazon Web Services (March 2019, expired March 2022)
Education
- Cybersecurity, BrainStation (2023)
Version 0.0.3 · PDF · Word · Markdown