Start Your Search Here

push notification bell

Would you like to receive notifications about jobs in Dallas?

push notification bell

You have blocked notifications

Oops! You have blocked notifications. Click here for more info

You have blocked notifications, please check your browser settings.

push notification bell

You're currently subscribed to job notifications

Want to change your notifications for job alerts?

push notification bell

Subscribe to notifications

You will no longer receive notifications

Job Search

Careers Integrated Resources Inc

Dallas / Global

Sr Cloud Reliability Engineer

Job Description

Sr. Cloud Reliability EngineerLocation: Dallas, TX 75201 Duration: 18 MonthsAbout the OpportunityAs a Senior Cloud Engineer in the Cloud SRE team, you will be responsible for designing and developing cloud solutions and engineering reliability tools for the Cloud Foundation Services (CFS) platform in the Infrastructure, Platforms & Operations organization. You will apply software engineering practices to build scalable, reusable solutions and utilities that enhance platform reliability across the client.QualificationsBachelor's degree in computer science, Information Systems, or equivalent background or equivalent experience7+ years of extensive experience in software development with focus on reliability and platform engineering5+ Years of advanced Python development skills with proven experience building enterprise-grade, highly available tools, APIs, and utilities3+ years of hands-on experience developing solutions in AWS environments with deep understanding of core services (EC2, VPC, S3, Lambda, IAM, CloudFormation, EventBridge, Step Functions etc.) and resource cost optimization3+ years of experience applying SRE principles including observability, toil automation, SLIs/SLOs and reliability engineeringExpert-level proficiency with Infrastructure as Code (IaC) using Terraform, including module development and state managementStrong experience with CI/CD pipelines, automated testing frameworks, and DevOps practicesExperience with observability tools and practices including Grafana, AWS CloudWatch, AWS CanaryExperience defining, implementing, and managing SLOs/SLIs and error budgets; familiarity with conducting RCAs and producing postmortem documentationWorking experience in Agile and Scaled Agile environments and familiarity with ITSM processes (incident, change, and problem management), resilience testing and chaos engineering practicesExperience with GoLang or additional programming languages is a plusResponsibilities: What Will Be Expected of YouDesign, develop, and maintain reliability solutions and SRE utilities to reduce toil, improve cloud platform reliability, and industrialize SRE practices across the systemBuild and optimize Infrastructure as Code (IaC) using Terraform to manage AWS resources related to SRE solutions, incorporating cost-efficient design principlesDevelop CI/CD pipelines and automated testing to ensure code quality, reliability, and rapid delivery of the solutionsDefine SRE standards, best practices, and guidelines for adoption across teams; establish SRE metrics like SLI, SLOs, etc.Apply software engineering best practices including version control, code reviews, test-driven development, and documentation to all developmentParticipate in incident management and on-call rotation, providing technical support for SRE tools, troubleshooting production issues, and collaborating with teams to reduce incident recurrence through proactive detection and pattern analysisStay current with emerging AWS services, SRE methodologies, and cloud-native development technologies, and drive adoption of innovative solutionsCollaborate within Agile and Scaled Agile frameworks with cross-functional teams to deliver integrated cloud automation solutionsProduce clear, blameless postmortems with actionable items and documented failure scenarios
Apply Now

Similar Opportunities

View all jobs

Get Job Alerts

Don't miss the perfect fit. Get Daily curated job alerts.

Job Title or Keyword(s)
Location