Start Your Search Here

push notification bell

Would you like to receive notifications about Computer and Mathematical Occupations jobs in Columbia?

push notification bell

You have blocked notifications

Oops! You have blocked notifications. Click here for more info

You have blocked notifications, please check your browser settings.

push notification bell

You're currently subscribed to job notifications

Want to change your notifications for job alerts?

push notification bell

Subscribe to notifications

You will no longer receive notifications

Job Search

Jobtailor

Columbia / Global

Data Engineer

Job Description

Design, develop, ingest, and maintain well-architected data pipelines that retrieve data from external feeds (APIs, SFTP, HTTPS, FTP, web scraping, Direct Connect) and internal agency sources into the data lake landing zone and downstream curated zones. Develop production-grade ETL workflows using AWS Glue, PySpark, Python, Lambda, and EMR, integrated with a shared ETL common library and orchestrated via Amazon Managed Workflows for Apache Airflow (MWAA). Load data accurately and optimally into S3 zones (Parquet, ORC, Iceberg), relational datastores (PostgreSQL, Redshift, Oracle), NoSQL databases, and knowledge bases/vector stores, preventing duplicate loads and maintaining data integrity and traceability across all lifecycle stages. Implement schema enforcement, XSD validation, data quality checks, error handling, and automated SNS notifications; ensure all production jobs populate ETL Load Reports and Gap Reports through static and dynamic ETL metadata. Develop semantic-layer objects (tables, views, materialized views) that ensure complete data coverage, optimized query performance, and consistent application of business logic. Develop XML parsing/shredding logic for high-volume regulatory filings using Glue PySpark, supporting schema evolution and batch processing per program standards. Design pipelines with query performance in mind and support rollback, reload, and date-range reprocessing capabilities without manual intervention. Support self-service ETL development by other agency teams through standardized, reusable components aligned with program standards, and facilitate the transition of externally developed ETL jobs into the Data Engineering team's production support. Create and maintain required engineering artifacts, including business requirements, ETL design documents, mapping documents, data models, data dictionaries, deployment references, operations and maintenance guides, and test plans. Deploy code through automated CI/CD pipelines using CloudFormation templates, following agency release, security, and governance processes. Provide operational support for production jobs, including rapid identification and resolution of failed jobs and performance issues; participate in on-call/after-hours support for production outages and emergencies as part of a team rotation. Collaborate with Data Officers, Data Stewards, SMEs, data providers, and IV&V teams to understand requirements and deliver user-accepted solutions; engage closely with the Product Owner and cross-functional teams to provide timely updates and resolve issues. Leverage AI-assisted development tools to accelerate coding, optimize workflows, and enhance code quality while adhering to security and performance standards. Work in Agile teams; drive iterative delivery, joint problem-solving, and continuous improvement, including participation in sprint planning and program increment planning. Requirements Bachelor's degree in Data Science, Computer Science, Engineering, or a related technical discipline. In lieu of a degree, four additional years of related experience is required Minimum of 5 years of related data engineering experience developing and deploying data pipelines in production. Hands-on experience developing ETL pipelines in AWS using Glue, Spark/PySpark, Lambda, and S3. Proficiency in Python (adhering to PEP 8) and strong SQL skills, with experience integrating data from relational databases. Experience processing structured and semi-structured data formats, including JSON, XML, CSV, Avro, and Parquet. Experience with workflow orchestration (Airflow/MWAA preferred). Experience with relational databases (PostgreSQL, Redshift, or Oracle) and lake table/file formats (Iceberg, Parquet, ORC). Experience with Git/GitHub version control and Agile methodologies. Strong attention to detail with a commitment to delivering high-quality and accurate work; excellent written and verbal communication skills. Ability to obtain and maintain a Moderate Risk Public Trust clearance; residing in the United States Core Competencies Demonstrates expertise in designing and developing data pipelines using AWS Glue, PySpark, and Python, with a strong focus on ETL workflows and data integrity. Proficient in collaborating with cross-functional teams to deliver high-quality data solutions while adhering to Agile methodologies. Highest-signal resume keywords AWS Glue ETL Development PySpark Programming Data Pipeline Design SQL Proficiency Agile Methodologies ATS Optimization Keywords Hard Skills ETL Development Data Pipeline Architecture Python Programming SQL Skills Data Quality Checks Schema Enforcement XML Parsing Workflow Orchestration Data Lake Management Version Control (Git/GitHub) Soft Skills Attention to Detail Written Communication Verbal Communication Collaboration Problem-Solving Industry Keywords Data Engineering Data Quality Data Integrity Regulatory Filings Agile Development Tools & Technologies Amazon S3 Amazon EMR Amazon Managed Workflows for Apache Airflow (MWAA) CloudFormation Data Lake Formats (Parquet, ORC, Iceberg) #J-18808-Ljbffr

Apply Now

Similar Opportunities

View all jobs

Get Job Alerts

Don't miss the perfect fit. Get Daily curated job alerts.

Job Title or Keyword(s)
Location