Npv
Data Engineer
$90,000 to $150,000 a year
About the job
tl;dr: Data Engineer; post-acquisition profitable health/adtech; processing 20TB+ daily at near-realtime speed; Python, Spark; remote, up to 150k USD base, we can talk higher figures and EU/UK/US employment
PulsePoint sits at the intersection of healthcare and adtech. We help brands and agencies interpret the hard-to-read signals across the health journey and unify these digital determinants of health with real-world data to produce the most dimensional view of the customer. We are 300+ and growing, post-acquisition business and one of the leading players in the US healthcare ad market.
We're looking for a Data Engineer to join a team, playing a key role at a technology company experiencing exponential growth. The data pipeline processes over 80 billion impressions a day (20TB+ of data, 220TB uncompressed), powering reports, budget updates, and optimization engines against extremely tight SLAs, with stats and reports delivered as close to real-time as possible.
Team responsibilities
- Install, maintain, and monitor Kafka, Hadoop, Presto, and RDBMS systems.
- Ingest, validate, and process internal and third-party data.
- Create, maintain, and monitor data flows in Hive, SQL, and Presto for consistency, accuracy, and lag time.
- Maintain and enhance the framework for jobs (primarily aggregate jobs in Hive).
- Build Kafka consumers using Spark Streaming for near-real-time aggregation.
- Train developers and analysts on data tools.
- Evaluate, select, and implement new tools.
- Handle backups, retention, high availability, and capacity planning.
- Review and approve DDL for databases, Hive framework jobs, and Spark Streaming to ensure standards are met.
- Participate in 24×7 on-call rotation for production support.
Stack
- Airflow, Docker, Graphite/Beacon, Hive, Impala, Kafka, Kubernetes, Presto, Spark Streaming, SQL Server, Sqoop.
Requirements
- BA/BS degree in Computer Science or a related field.
- 5+ years of software engineering experience.
- Strong Spark expertise, Spark Streaming is highly desirable.
- Proficiency in Linux.
- Fluency in Python; experience with Scala/Java is a strong plus.
- Strong understanding of RDBMS and SQL.
- Willingness to participate in 24×7 on-call rotation.
Nice to have
- Knowledge and exposure to distributed production systems like Hadoop.
- Knowledge and exposure to cloud migration.
We offer
- Remote work, high engineering bar and comfortable culture.
- Flat hierarchy with easy access to business, product, and operations.
- Enormous scale (80B+ impressions/day) with real growth potential.
- Ownership and direct impact, you have room to shift focus as your interests evolve.
- Up to 150k USD salary, higher figures and EU/UK/US employment are negotiable.