Resume

Pavan Badempet - Data Engineer

Data Engineer with 3 years of experience building high-throughput distributed ETL data pipelines, Medallion Lakehouse platforms, and on-prem/cloud infrastructure across BFSI (Nomura Capital) and Automotive (Nissan) domains. Proficient in Python, SQL and PySpark performance tuning, dimensional modeling, workflow orchestration using Airflow and AutoSys, and data platform modernization with Databricks, AWS and Gen AI Integrated Development.

What I Do
Distributed Data Pipelines
Architecting scalable batch extraction and transformation jobs using Spark, Spark SQL, and Databricks inside Apache Iceberg and Delta Lake Medallion Lakehouses.
Cloud & Event-Driven Workflows
Developing event-driven serverless ingestion solutions using AWS Lambda, Step Functions, EventBridge, and SNS, handling 100K+ events daily with idempotent delivery.
Spark Performance Tuning
Optimizing distributed computation via Adaptive Query Execution, broadcast joins, dynamic partition pruning, caching, and shuffle tuning, reducing runtime by 30%.
GenAI & Vector Lakehouse Integration
Deploying RAG pipelines, vector retrieval engines (FAISS, pgvector), and low-latency serving APIs using FastAPI, Docker, and Apache Airflow orchestration.
Experience
Nov 2023 – Present
Data Engineer - Tata Consultancy Services
  • Nomura Capital: Architected scalable batch extraction and transformation jobs using Spark and Spark SQL inside an Apache Iceberg Medallion Lakehouse at Nomura Capital, handling 1B+ records per month and serving 200K+ analytical queries daily.

  • Optimized distributed computation via Adaptive Query Execution, broadcast joins, dynamic partition pruning, caching, and shuffle tuning, reducing SLA-bound runtime by 30%.

  • Migrated workloads from legacy YARN and HDFS to Kubernetes with MinIO object storage, resolving dependency isolation challenges and lowering infrastructure spend by 25%.

  • Spearheaded automation of orchestration schedules and deployment configurations, maintaining 99.9% on-time completion rate across 50+ recurring batch runs.

  • Diagnosed failures via executor logs and DAG execution plans, constructed automated retries, fallback recovery handlers, and operational runbooks, cutting Mean Time to Recovery (MTTR) by 30%.

  • Nissan: Developed event-driven cloud capture solutions leveraging AWS Lambda, Step Functions, and EventBridge at Nissan, handling 100K+ events daily with idempotent delivery and schema validation via a Streamlit dashboard.

  • Established programmatic quality validation rules and version-controlled table evolution handlers across 20+ source systems, improving output reliability to 99.5% for downstream consumers.

Education
Aug 2019 – Sep 2023
Institution: Guru Nanak Institutions Technical Campus

Degree: Bachelor of Technology in Computer Science and Engineering
Minor: AI/ML | Grade: CGPA 8.2/10 (First Class with Distinction)
Minor Grade: CGPA 8.6/10
Scholarship: Prime Minister’s Scholarship (PMSS)

Languages
  • Python
  • SQL
  • Scala
  • Java
Big Data & Lakehouse
  • Apache Spark (PySpark, Spark SQL)
  • Databricks & Unity Catalog
  • Structured Streaming
  • Apache Iceberg & Delta Lake
  • Medallion Architecture
  • Dremio, MinIO, HDFS
Modeling & Quality
  • Dimensional Modeling (SCD Type 2, Star Schema)
  • CDC & Schema Evolution
  • Data Quality & Validation
Cloud & Tools
  • AWS (S3, Lambda, Step Functions, EventBridge, SNS)
  • Apache Airflow & AutoSys
  • Docker & Git / GitHub Actions
  • PostgreSQL & SQL Server
  • FastAPI & Streamlit
Agent View