OpenBrand
Chicago / Global
You have blocked notifications
Oops! You have blocked notifications. Click here for more info
You have blocked notifications, please check your browser settings.
You're currently subscribed to job notifications
Subscribe to notifications
You will no longer receive notifications
Chicago / Global
Senior DevOps / Cloud Infrastructure Engineer Location: Hybrid (Chicago, IL)
Be a part of a fast-growing, winning team helping Fortune 1000 consumer brands and retailers leverage AI-driven data insights.
OpenBrand is one of the world's most respected market intelligence companies. OpenBrand's data and market research products give manufacturers, retailers, and industry players a competitive edge across a wide range of industries (including IT, consumer electronics, home appliances, health, wellness, beauty, small appliances, and other consumer durables) and help marketing, product, sales, and pricing teams make more informed decisions in a rapidly changing market environment.
This is a permanent, hands-on Senior DevOps / Cloud Infrastructure Engineer role owning the cloud infrastructure, delivery pipelines, and operational reliability behind OpenBrand's production data platform — the AWS estate, CI/CD and package tooling, observability, and incident response.
You will work at the intersection of cloud architecture, automation, and operational excellence, with a strong emphasis on:
Owning the AWS estate end to end — architecture, cost, performance, and resilience
Building and maintaining reliable CI/CD pipelines, package dependencies, and deployment automation
Establishing real operating control — deploy, observe, troubleshoot, and recover — across the full platform
Executing infrastructure migrations cleanly, on schedule, and without customer or data disruption
This is a build-and-run role, not a purely advisory one. You will own infrastructure end to end — design it, automate it, deploy it, monitor it, and improve it — with real accountability for uptime, cost, and delivery velocity.
Your first year is an integration year. OpenBrand has grown through acquisition, and the near-term priority is consolidating inherited platform infrastructure onto OpenBrand standards — separating cloud accounts, proving out monitoring and recovery, and retiring legacy dependencies. Success here means being effective inside imperfect inherited systems: making production safe and observable first, then modernizing deliberately.
What comes after is the larger half of the job. Once consolidation is complete, this role owns the platform's forward roadmap — cloud cost re-architecture, moving workloads onto managed and containerized services, deepening automation and observability, and scaling the infrastructure behind a growing data business. We are hiring an owner for the platform, not a migration.
You will collaborate closely with Engineering, Data, Security, and IT teams, reporting to the VP of DevOps. This role does not focus on people management, but requires strong judgment, independence, and end-to-end ownership.
Key Responsibilities:
Cloud Infrastructure & Cost Optimization Own the AWS environment across production, QA, and shared services accounts — including account strategy, IAM, networking, EKS/ECS/Lambda, S3, logging, and billing structure
Lead right-sizing and re-architecture of a large EC2 footprint, moving workloads to appropriate instance families, purchase models, and managed services
Build and maintain infrastructure as code (Terraform, CloudFormation, or equivalent) so environments are reproducible and reviewable
Own DNS, certificates, and CDN configuration, including renewal automation and domain transitions
Establish cost visibility — tagging standards, allocation reporting, budgets, and anomaly alerting — and drive measurable reductions in cloud spend
Plan and execute data center and legacy workload decommissioning, including migration sequencing and rollback planning
Design for resilience: multi-AZ posture, documented recovery objectives, and tested backup restores — validated by actual recovery, not by the presence of a backup job
Own capacity planning and the infrastructure roadmap as the business grows — evaluating managed services, containerization, and architectural changes on their operational and cost merits
CI/CD, Observability & Platform Reliability Design, maintain, and improve CI/CD pipelines (GitLab CI, GitHub Actions, or equivalent) from commit through production deployment, including runner fleets and build environments
Containerize and orchestrate services; manage image registries, build caching, and artifact promotion across environments
Own package and artifact dependencies — ECR, JFrog, npm, PyPI, and private mirrors — so builds are reproducible and not silently dependent on external or third-party infrastructure
Implement observability — metrics, logging, tracing, dashboards, and actionable alerting (Grafana, Sentry, CloudWatch, or similar) — so failures are detected before customers see them
Own the incident-response path: alert routing and escalation (Opsgenie, PagerDuty, or similar), on-call rotation, P1/P2 severity definitions, and post-incident review
Write and test runbooks for critical services, so restart and troubleshooting steps are proven rather than assumed
Automate away manual toil: provisioning, patching, certificate rotation, secrets distribution, and routine operational tasks
Manage secrets and shared credentials through proper tooling (AWS Secrets Manager, vault-style platforms) rather than ad-hoc practice, including key rotation and access removal after workload transfer
Integration, Consolidation & Cross-Functional Collaboration Execute infrastructure separation and integration work arising from acquisitions — account transfers, network segmentation, identity migration, and consolidation onto OpenBrand standards
Build the infrastructure dependency map: what must run on day one, what is shared, what migrates, what gets replaced, and what retires — with owners and target dates
Produce objective evidence of operating independence: validated deploys, working monitoring, tested restores, and an exercised incident path — not just credentialed access
Partner with Data Engineering to keep production data pipelines and analytical platforms (Snowflake, OpenSearch, Airflow, or equivalent) reliable through periods of change, where infrastructure touches data movement
Partner with Security and Compliance on IAM design, access reviews, endpoint and network posture, and SOC 2 evidence collection
Document environments, runbooks, and topology so operational knowledge does not sit with one person
Support Engineering and Product teams with infrastructure input to roadmap decisions and delivery planning
Communicate infrastructure risk clearly to non-technical stakeholders, and help establish best practices for change management, reproducibility, and operational handoff
Qualifications:
Education & Experience 5+ years of experience in DevOps, Site Reliability, Cloud Operations, or Platform Engineering, with significant hands-on responsibility for production infrastructure
BS degree in a technical field (e.g., Computer Science, Computer Engineering, Information Systems) or equivalent practical experience
Technical Skills Deep, hands-on AWS production operations experience across IAM, VPC/networking, EKS/ECS/Lambda, EC2, S3, RDS, logging, backup/DR, and account strategy
Strong proficiency with infrastructure as code (Terraform preferred) and Kubernetes deployment models (Helm, jsonnet, or equivalent)
Advanced proficiency in Linux administration and scripting (Python, Bash, or equivalent)
Demonstrated experience building and operating CI/CD pipelines for containerized services, including runners, Docker/ECR, and private package management (JFrog, npm, PyPI, artifact mirroring)
Practical experience with monitoring and observability platforms (Grafana, Prometheus, Sentry, CloudWatch, Datadog, or similar), plus real on-call and incident-response experience — treating monitoring as operational response, not dashboard creation
Working knowledge of networking fundamentals — routing, VPN, DNS, certificates, firewall rules, and network segmentation
Security hygiene: secrets management, key rotation, root/admin ownership, least privilege, and access removal
Proven ability to stabilize and operate inherited or legacy production systems while modernizing them incrementally
Ability to write clear, usable runbooks and communicate technical risk to non-technical stakeholders
Experience using Large Language Models (LLMs) (e.g., GPT-based or similar) to improve productivity and efficiency in engineering workflows, including tasks such as scripting, infrastructure code review, documentation, troubleshooting, and runbook development
Experience designing workflows that are robust, testable, and maintainable over time
Comfort working across the full lifecycle: design → automation → deployment → monitoring → iteration
Preferred
Chicago / Global
Chicago / Global
Chicago / Global
Chicago / Global
Chicago / Global
Chicago / Global