Staff Data Engineer
About the Role
Intellias is looking for a Staff Data Engineer to join a large-scale digital healthcare initiative focused on building the data and AI foundations that power personalized healthcare experiences.
This is a Staff-level Individual Contributor role for an engineer who combines deep data engineering expertise with distributed systems thinking, backend development, and technical leadership. You will design and build scalable data infrastructure, solve complex cross-platform engineering challenges, and establish technical patterns and standards adopted across multiple engineering teams.
The role is highly hands-on while also requiring the ability to provide technical direction, mentor experienced engineers, and influence architecture across the broader platform.
Project Overview
The project is building a FHIR-based healthcare data platform that integrates data from EHRs, healthcare applications, wearables, portals, and other sources to enable connected and personalized healthcare experiences.
The platform supports:
- Real-time and batch data processing
- Operational and analytical workloads
- Healthcare interoperability
- Data-intensive product capabilities
- Machine learning and AI use cases
- Scalable, cloud-native data infrastructure
The technology landscape includes Python, PySpark, Apache Spark, Kafka, DuckDB, Prefect/Airflow, FastAPI, MongoDB, Docker, Kubernetes, and cloud-native infrastructure.
The core engineering challenge is to build secure, reliable, and scalable data foundations capable of processing heterogeneous healthcare data while supporting product engineering, analytics, machine learning, and emerging AI capabilities.
What You'll Do
- Design and implement solutions for large-scale, ambiguous data engineering and distributed-systems challenges.
- Architect, build, and evolve real-time and batch data infrastructure supporting operational, analytical, and AI/ML workloads.
- Design and scale event-driven data pipelines using technologies such as Kafka, Spark, and PySpark.
- Build production backend services and APIs using Python and FastAPI to support data and platform capabilities.
- Define scalable approaches to data ingestion, transformation, modeling, validation, storage, and delivery.
- Design data architectures that address schema evolution, data quality, heterogeneous data integration, scalability, and operational reliability.
- Establish and champion technical standards, architectural patterns, engineering best practices, and reusable platform components.
- Serve as a cross-team technical authority for data systems, distributed processing, pipelines, and platform architecture.
- Improve the reliability, scalability, performance, and cost efficiency of data infrastructure.
- Implement robust observability, logging, monitoring, and alerting across distributed data systems.
- Design systems with security, privacy, healthcare compliance, and scalability built in from the beginning.
- Improve engineering workflows, CI/CD pipelines, automated testing, and production deployment practices.
- Collaborate closely with Product, ML/AI, Platform, Infrastructure, and application engineering teams to develop shared data capabilities.
- Provide technical mentorship through architecture reviews, design discussions, pairing, and knowledge sharing.
- Evaluate emerging data and AI technologies and introduce them where they provide measurable improvements to platform capabilities or engineering efficiency.
- Remain hands-on with production engineering, validating architectural decisions through implementation and operational experience.
What You Bring
Required Qualifications
- 8+ years of professional software and data engineering experience, including at least 3 years working with large-scale distributed data systems.
- Deep expertise designing and building scalable data platforms, pipelines, and distributed processing systems.
- Strong hands-on proficiency in Python, with production experience using technologies and libraries such as PySpark, Pandas, and FastAPI.
- Strong experience with Apache Spark/PySpark and distributed data processing at scale.
- Experience designing and implementing real-time and batch data pipelines using Kafka, Spark, or comparable distributed processing technologies.
- Strong understanding of data architecture, data modeling, schema evolution, data quality, and heterogeneous data integration.
- Experience with workflow orchestration technologies such as Prefect, Apache Airflow, or equivalent.
- Experience building production backend services and APIs supporting data and platform capabilities.
- Strong SQL skills and hands-on experience with modern data storage technologies, including relational and NoSQL databases such as MongoDB.
- Strong understanding of distributed systems principles, including scalability, reliability, fault tolerance, consistency, and performance.
- Experience designing and implementing observability, logging, monitoring, and alerting for distributed data systems.
- Strong cloud-native engineering experience, including Docker, Kubernetes, CI/CD, and production deployment practices.
- Proven ability to independently design solutions for ambiguous, large-scale engineering problems spanning multiple systems and teams.
- Demonstrated experience establishing technical standards, engineering patterns, development practices, and reusable platform capabilities.
- Strong technical leadership skills, including experience mentoring senior engineers, conducting design reviews, and influencing architecture without formal people-management authority.
- Strong understanding of security, privacy, and data protection principles for production data platforms.
- Professional English proficiency sufficient for direct collaboration with U.S.-based engineering, product, ML, and platform teams.
- Bachelor's degree in computer science, Engineering, or a related field.
Nice to Have
- Experience with FHIR, HL7, C-CDA, or other healthcare interoperability standards.
- Previous experience in healthcare, HealthTech, or another regulated data environment.
- Understanding of HIPAA, HITECH, or comparable data privacy and compliance requirements.
- Experience deploying or supporting LLMs, ML models, AI-enabled data products, or retrieval-based systems.
- Experience building data infrastructure supporting ML/AI workloads and production inference.
- Hands-on experience with observability technologies such as OpenTelemetry, Datadog, Prometheus, or similar platforms.
- Experience with DuckDB or modern analytical data-processing technologies.
- Experience with AI-assisted engineering tools such as Claude Code, GitHub Copilot, Cursor, or comparable solutions.
- Contributions to open-source projects or publicly available engineering initiatives.
Why Join This Initiative
This role offers the opportunity to shape the data and AI foundations of a modern healthcare platform operating at significant scale.
You will work on technically complex challenges spanning:
- Distributed data processing
- Event-driven architectures
- Backend services
- Cloud-native infrastructure
- Healthcare interoperability
- Data platform architecture
- AI/ML enablement
As a Staff-level Individual Contributor, your impact will extend well beyond individual services or pipelines. You will establish architectural patterns, influence technical decisions across teams, mentor experienced engineers, and help define how the broader engineering organization builds and operates data-intensive systems.
At the same time, this is fundamentally a hands-on engineering role. You will design systems, build critical components, validate architectural decisions through code, and remain closely connected to how systems behave in production.
This position is particularly well suited for an engineer who wants to combine deep technical expertise with organization-wide engineering influence while building data infrastructure that directly enables better healthcare products and emerging AI capabilities.