Scorpion Therapeutics
Indianapolis / Global
Data Engineer
- $120.000 - $180.000
You have blocked notifications
Oops! You have blocked notifications. Click here for more info
You have blocked notifications, please check your browser settings.
You're currently subscribed to job notifications
Subscribe to notifications
You will no longer receive notifications
Indianapolis / Global
What You Will Do Data Engineering & Pipeline Development Design, develop, and optimize scalable data pipelines using Databricks, PySpark, Python, SQL, and Delta Lake to ingest, transform, and load data into data warehouses and data lakes.
Build Databricks pipelines implementing canonical data models across the medallion architecture (Bronze → Silver → Gold) within the CE trust boundary.
Evaluate/apply Databricks capabilities (Unity Catalog, Delta Lake, Databricks Workflows, serverless compute, Lakebase, ingestion connectors) based on performance, cost, and scalability.
Implement and maintain ELT/ETL workflows (Databricks Workflows, Auto Loader, Structured Streaming, Delta Live Tables).
Build and maintain CI/CD pipelines (GitHub Actions; dev → test → prod promotion) for CE data and artifacts.
Automate data ingestion and product creation to reduce manual maintenance and onboarding.
Data Governance, Quality & Security Implement data governance policies ensuring data quality, integrity, security, and compliance (e.g., GxP, HIPAA) and covered-entity constructs.
Implement row/column-level security, masking, and tokenization to enforce PHI isolation.
Use Unity Catalog for metadata, lineage, and access control.
Establish testing/validation (pytest, DLT/Great Expectations) and monitoring/alerting.
Data Modeling & Architecture Develop/maintain data models, schemas, metadata; follow Lakehouse/Medallion principles.
Create reusable transformation frameworks and automated data quality checks.
Partner on reference architecture; document pipelines/processes.
Collaboration & Innovation Translate stakeholder data requirements into technical solutions.
Monitor performance, troubleshoot, and improve availability/reliability.
Participate in code reviews/architecture discussions; promote best practices.
Evaluate/recommend new tools; support production integration of ML/AI tools with PHI classification/consent.
Participate in Agile ceremonies (Jira or equivalent).
Your Minimum Basic Qualifications Bachelor’s degree in CS/Engineering/IS or related quantitative field.
5+ years data engineering/ETL development.
Proficient in SQL and at least one language (e.g., Python/Java/Databricks).
Experience with cloud data platforms (Databricks, AWS/Azure/GCP) and services (S3, Redshift, Snowflake, ADLS, BigQuery).
Proficient with Git-based CI/CD workflows (GitHub Actions or equivalent).
What You Should Bring Excellent problem-solving; ability to build/test pipelines from architecture.
Strong communication/collaboration.
Prior pharma/life sciences experience (preferred).
#J-18808-Ljbffr
Indianapolis / Global
Indianapolis / Global
New Whiteland / Global
Indianapolis / Global
Indianapolis / Global
Indianapolis / Global