// senior data engineer @ disney streaming

Saiteja Desu

I build the data platforms that big decisions ride on — moving petabytes across streaming media and semiconductor manufacturing, and turning fragile manual data operations into automated, observable systems.

10+ years in data engineering
PB scale platforms operated
23 countries, unified reporting
CEO award, GlobalFoundries ’20
career.dag operational · records processed: 8,431,204,117

↳ this page runs like my pipelines — pick a stage to jump in. scroll the diagram sideways on small screens.

01about

For over a decade I’ve done one thing well: make large-scale data infrastructure trustworthy. Pipelines that don’t silently break, backfills that don’t need a human babysitter, metadata that analytics and recommendation systems can actually rely on.

At Disney Streaming I led the integration of Hotstar’s data into the Disney ecosystem — unlocking unified content-performance reporting across ~23 countries — and authored a config-driven backfill framework my teammates now reuse well beyond its original purpose. Before that, at GLOBALFOUNDRIES, I was a key contributor to the cloud-native ingestion platform behind the company’s global data lake — thousands of streaming pipelines carrying petabyte-scale manufacturing data — and built its monitoring application from scratch, work recognized with the company’s 2020 CEO Award.

These days I’m also giving back upstream: contributing code to Apache Airflow and open energy-data projects like PUDL and OED.

02experience

Disney Streaming

Senior Data Engineer · Jun 2022 – present · SF Bay Area

  • Led the integration of Hotstar data into Disney’s ecosystem across ~23 countries, solving deep metadata-matching challenges to unlock unified content-performance reporting for EMEA and APAC markets.
  • Architected a config-driven backfill framework that turned static SQL backfills into dynamic, parameterized workflows — now adopted by engineers across the team and cutting reprocessing effort and cost dramatically.
  • Design and operate Databricks / PySpark pipelines that feed analytics and recommendation engines, optimizing compute cost while improving reliability and runtime.
  • Built and operationalized the content-metadata systems behind launch-critical reporting for major product rollouts; the team’s go-to expert for metadata and pipeline architecture.
  • Mentor junior engineers on delivery quality and technical ownership.

GLOBALFOUNDRIES

Data Engineer · Jul 2019 – May 2022 · Malta, NY

  • Key contributor to the cloud-native ingestion platform that became the foundation of the company’s global data lake — thousands of streaming pipelines, petabyte-scale manufacturing data on AWS (EMR, Glue, S3, Lambda, Redshift).
  • Designed and built a pipeline-health monitoring application from scratch, shifting the team from reactive firefighting to proactive issue detection across thousands of production pipelines.
  • Partnered with AI/ML teams to deliver the high-fidelity data behind predictive yield models for semiconductor manufacturing.
  • Improved data-mart performance and scalability by eliminating bottleneck queries in Oracle environments.
  • Recognized with the 2020 CEO Award and multiple Spotlight Awards for innovation and intelligent automation.

San Jose State University

Business Intelligence Intern · May 2018 – May 2019 · San Jose, CA

  • Delivered data-driven insights to department chairs to support student success, via interactive OBIEE dashboards.
  • Developed complex ETL jobs with PL/SQL and collaborated on star- and snowflake-schema data models.

Tata Consultancy Services

Systems Engineer · Jun 2014 – Jul 2017 · Mumbai, India

  • Delivered $120K in annual cost savings by designing reusable components and automating data workflows.
  • Cut data-warehouse load time by 40% by optimizing Informatica transformations and shell scripts.
  • Recognized with “Star of the Month” and “On the Spot” awards during critical project periods.

03selected work

hackathon · top 7

CareGraph

AI-powered senior-care intelligence: daily voice check-ins, Neo4j graph reasoning, and AI-generated care recommendations that surface health risks early. Placed in the top 7 with strong reviews from judges and sponsors.

Neo4j Aura · FastAPI · CrewAI · Qwen3-235B · Bland AI

disney streaming

Config-Driven Backfill Framework

Turned static SQL backfill scripts into dynamic, parameter-driven workflows for large-scale historical reprocessing. Adopted by engineers across the team — the pattern outgrew its original use case.

SQL · Databricks · ETL orchestration · pipeline design

disney streaming

Hotstar → Disney Data Integration

Unified an acquired streaming service’s data into the parent ecosystem using layered fuzzy metadata matching — enabling one consistent view of content performance across ~23 countries for global stakeholders.

PySpark · Databricks · cross-platform data modeling

globalfoundries

Pipeline Health Monitoring Platform

A monitoring application built from the ground up to watch thousands of production pipelines feeding a petabyte-scale data lake — replacing reactive troubleshooting with proactive detection for manufacturing-critical data.

PySpark · Java · AWS Lambda / EMR / S3 / Glue · Redshift

04open source

Apache Airflow

merged ✓

Code contributor to the orchestration system that runs most of the world’s production data platforms. First substantive pull request merged into main after committer review and milestoned for the Airflow 3.2.2 release — with more contributions in flight across scheduler and operator behavior.

PUDL & Open Energy Dashboard

energy data

Contributor to open energy-data projects — the Public Utility Data Liberation project and OED — helping make public infrastructure data usable for researchers and analysts.

05toolbox

languages

PythonSQLPySparkJavaScalaPL/SQLR

cloud & big data

DatabricksDelta LakeApache SparkAWS · EMR / S3 / Glue / LambdaRedshiftSnowflakeGCPDockerKubernetes

data engineering

Apache AirflowETL designBackfill automationDimensional modelingInformaticaPipeline observability

ai & analytics

Claude & agentic workflowsCursorNeo4jScikit-LearnPandasTableauPower BI

06recognition

  • 2020CEO Award — GLOBALFOUNDRIES
  • 19–22Multiple Spotlight Awards — GLOBALFOUNDRIES
  • 14–17“Star of the Month” & “On the Spot” — TCS
  • certDatabricks Certified Associate Developer — Apache Spark 3.0
  • 2026Top-7 finish — CareGraph, AI hackathon

07education

  • San Jose State University

    M.S. Computer Software Engineering — Cloud & Data Science · 2017–2019

  • KL University

    B.Tech. Computer Science & Engineering · 2010–2014

08sink — where the data lands

Always happy to talk data, pipelines, or interesting problems. Say hi.