Start Your Search Here

push notification bell

Would you like to receive notifications about Computer and Mathematical Occupations jobs in Chicago?

push notification bell

You have blocked notifications

Oops! You have blocked notifications. Click here for more info

You have blocked notifications, please check your browser settings.

push notification bell

You're currently subscribed to job notifications

Want to change your notifications for job alerts?

push notification bell

Subscribe to notifications

You will no longer receive notifications

Job Search

OpenBrand

Chicago / Global

Senior DevOps / Cloud Infrastructure Engineer

Job Description

Senior DevOps / Cloud Infrastructure Engineer Location: Hybrid (Chicago, IL)

Be a part of a fast-growing, winning team helping Fortune 1000 consumer brands and retailers leverage AI-driven data insights.

OpenBrand is one of the world's most respected market intelligence companies. OpenBrand's data and market research products give manufacturers, retailers, and industry players a competitive edge across a wide range of industries (including IT, consumer electronics, home appliances, health, wellness, beauty, small appliances, and other consumer durables) and help marketing, product, sales, and pricing teams make more informed decisions in a rapidly changing market environment.

This is a permanent, hands-on Senior DevOps / Cloud Infrastructure Engineer role owning the cloud infrastructure, delivery pipelines, and operational reliability behind OpenBrand's production data platform — the AWS estate, CI/CD and package tooling, observability, and incident response.

You will work at the intersection of cloud architecture, automation, and operational excellence, with a strong emphasis on:

Owning the AWS estate end to end — architecture, cost, performance, and resilience

Building and maintaining reliable CI/CD pipelines, package dependencies, and deployment automation

Establishing real operating control — deploy, observe, troubleshoot, and recover — across the full platform

Executing infrastructure migrations cleanly, on schedule, and without customer or data disruption

This is a build-and-run role, not a purely advisory one. You will own infrastructure end to end — design it, automate it, deploy it, monitor it, and improve it — with real accountability for uptime, cost, and delivery velocity.

Your first year is an integration year. OpenBrand has grown through acquisition, and the near-term priority is consolidating inherited platform infrastructure onto OpenBrand standards — separating cloud accounts, proving out monitoring and recovery, and retiring legacy dependencies. Success here means being effective inside imperfect inherited systems: making production safe and observable first, then modernizing deliberately.

What comes after is the larger half of the job. Once consolidation is complete, this role owns the platform's forward roadmap — cloud cost re-architecture, moving workloads onto managed and containerized services, deepening automation and observability, and scaling the infrastructure behind a growing data business. We are hiring an owner for the platform, not a migration.

You will collaborate closely with Engineering, Data, Security, and IT teams, reporting to the VP of DevOps. This role does not focus on people management, but requires strong judgment, independence, and end-to-end ownership.

Key Responsibilities:

Cloud Infrastructure & Cost Optimization Own the AWS environment across production, QA, and shared services accounts — including account strategy, IAM, networking, EKS/ECS/Lambda, S3, logging, and billing structure

Lead right-sizing and re-architecture of a large EC2 footprint, moving workloads to appropriate instance families, purchase models, and managed services

Build and maintain infrastructure as code (Terraform, CloudFormation, or equivalent) so environments are reproducible and reviewable

Own DNS, certificates, and CDN configuration, including renewal automation and domain transitions

Establish cost visibility — tagging standards, allocation reporting, budgets, and anomaly alerting — and drive measurable reductions in cloud spend

Plan and execute data center and legacy workload decommissioning, including migration sequencing and rollback planning

Design for resilience: multi-AZ posture, documented recovery objectives, and tested backup restores — validated by actual recovery, not by the presence of a backup job

Own capacity planning and the infrastructure roadmap as the business grows — evaluating managed services, containerization, and architectural changes on their operational and cost merits

CI/CD, Observability & Platform Reliability Design, maintain, and improve CI/CD pipelines (GitLab CI, GitHub Actions, or equivalent) from commit through production deployment, including runner fleets and build environments

Containerize and orchestrate services; manage image registries, build caching, and artifact promotion across environments

Own package and artifact dependencies — ECR, JFrog, npm, PyPI, and private mirrors — so builds are reproducible and not silently dependent on external or third-party infrastructure

Implement observability — metrics, logging, tracing, dashboards, and actionable alerting (Grafana, Sentry, CloudWatch, or similar) — so failures are detected before customers see them

Own the incident-response path: alert routing and escalation (Opsgenie, PagerDuty, or similar), on-call rotation, P1/P2 severity definitions, and post-incident review

Write and test runbooks for critical services, so restart and troubleshooting steps are proven rather than assumed

Automate away manual toil: provisioning, patching, certificate rotation, secrets distribution, and routine operational tasks

Manage secrets and shared credentials through proper tooling (AWS Secrets Manager, vault-style platforms) rather than ad-hoc practice, including key rotation and access removal after workload transfer

Integration, Consolidation & Cross-Functional Collaboration Execute infrastructure separation and integration work arising from acquisitions — account transfers, network segmentation, identity migration, and consolidation onto OpenBrand standards

Build the infrastructure dependency map: what must run on day one, what is shared, what migrates, what gets replaced, and what retires — with owners and target dates

Produce objective evidence of operating independence: validated deploys, working monitoring, tested restores, and an exercised incident path — not just credentialed access

Partner with Data Engineering to keep production data pipelines and analytical platforms (Snowflake, OpenSearch, Airflow, or equivalent) reliable through periods of change, where infrastructure touches data movement

Partner with Security and Compliance on IAM design, access reviews, endpoint and network posture, and SOC 2 evidence collection

Document environments, runbooks, and topology so operational knowledge does not sit with one person

Support Engineering and Product teams with infrastructure input to roadmap decisions and delivery planning

Communicate infrastructure risk clearly to non-technical stakeholders, and help establish best practices for change management, reproducibility, and operational handoff

Qualifications:

Education & Experience 5+ years of experience in DevOps, Site Reliability, Cloud Operations, or Platform Engineering, with significant hands-on responsibility for production infrastructure

BS degree in a technical field (e.g., Computer Science, Computer Engineering, Information Systems) or equivalent practical experience

Technical Skills Deep, hands-on AWS production operations experience across IAM, VPC/networking, EKS/ECS/Lambda, EC2, S3, RDS, logging, backup/DR, and account strategy

Strong proficiency with infrastructure as code (Terraform preferred) and Kubernetes deployment models (Helm, jsonnet, or equivalent)

Advanced proficiency in Linux administration and scripting (Python, Bash, or equivalent)

Demonstrated experience building and operating CI/CD pipelines for containerized services, including runners, Docker/ECR, and private package management (JFrog, npm, PyPI, artifact mirroring)

Practical experience with monitoring and observability platforms (Grafana, Prometheus, Sentry, CloudWatch, Datadog, or similar), plus real on-call and incident-response experience — treating monitoring as operational response, not dashboard creation

Working knowledge of networking fundamentals — routing, VPN, DNS, certificates, firewall rules, and network segmentation

Security hygiene: secrets management, key rotation, root/admin ownership, least privilege, and access removal

Proven ability to stabilize and operate inherited or legacy production systems while modernizing them incrementally

Ability to write clear, usable runbooks and communicate technical risk to non-technical stakeholders

Experience using Large Language Models (LLMs) (e.g., GPT-based or similar) to improve productivity and efficiency in engineering workflows, including tasks such as scripting, infrastructure code review, documentation, troubleshooting, and runbook development

Experience designing workflows that are robust, testable, and maintainable over time

Comfort working across the full lifecycle: design → automation → deployment → monitoring → iteration

Preferred

Apply Now

Get Job Alerts

Don't miss the perfect fit. Get Daily curated job alerts.

Job Title or Keyword(s)
Location