Data Engineer
ACT 01 // IDENTITY & ELEVATOR PITCH
00:00
Resume
Data Engineer // Portfolio Story

Pavan Badempet

Python, PySpark & SQL Specialist

Data Engineer with 3 years of experience at Tata Consultancy Services building high-throughput distributed ETL pipelines, Medallion Lakehouse platforms, and on-prem/cloud infrastructure across BFSI (Nomura Capital) and Automotive (Nissan) domains.

BFSI: Nomura Capital Automotive: Nissan Apache Spark / Databricks AWS Serverless GenAI & Vector RAG
Core Profile Hyderabad, India
3+ Years of Distributed Lakehouse & Cloud Engineering Experience
Distributed ETL
PySpark, Spark SQL, Databricks
Lakehouse Tables
Delta Lake, Apache Iceberg
Cloud Systems
AWS Lambda, Step Fn, EventBridge
AI & Serving
FastAPI, FAISS, Docker, RAG
Tata Consultancy Services // Client Engagement

Nomura Capital: High-Throughput Iceberg Medallion Lakehouse

Architected scalable batch extraction and transformation jobs using Spark and Spark SQL inside an Apache Iceberg Medallion Lakehouse powering mission-critical financial analytics.

1B+
Monthly Records Processed
High-volume batch pipelines processing distributed capital markets data into structured Lakehouse tables.
200K+
Daily Analytical Queries
Low-latency analytical queries served to downstream risk and reporting dashboards with zero bottlenecking.
99.9%
On-Time Pipeline SLA
Maintained strict completion rates across 50+ recurring automated batch jobs orchestrated via Airflow & AutoSys.
Apache Iceberg Medallion Architecture Flow (Nomura Capital)
Batch & Lakehouse Engine
BRONZE LAYER
Raw Ingestion
High-throughput raw data feeds ingested into partitioned Iceberg tables with audit metadata.
SILVER LAYER
Cleanse & Transform
PySpark cleaning, schema validation, deduplication, and SCD Type 2 dimension management.
GOLD LAYER
Curated Aggregates
Business dimensional models & financial rollups optimized for sub-second analytical reporting.
SERVING & QUERY
200K+ Queries/Day
High-concurrency query execution layer serving analysts and risk management systems.
Performance Engineering // Reliability & Cost

Spark Performance Tuning & Kubernetes Modernization

Slashing SLA runtimes by 30% and cutting infrastructure spend by 25% through deep distributed engine optimization and cloud-native migration.

-30%
SLA Runtime Reduction
Achieved via Adaptive Query Execution (AQE), broadcast joins, dynamic partition pruning, caching & shuffle tuning.
-25%
Infra Cost Reduction
Migrated legacy YARN & HDFS workloads to Kubernetes with MinIO object storage, ensuring clean dependency isolation.
-30%
Mean Time to Recovery (MTTR)
Diagnosed failures via DAG execution plans & executor logs; constructed automated retries and operational runbooks.

Spark Distributed Engine Optimizations

  • Adaptive Query Execution (AQE): Dynamically coalescing shuffle partitions and converting sort-merge joins to broadcast hash joins at runtime.
  • Broadcast Join Strategy: Broadcasting dimension lookups to eliminate expensive wide transformations and data shuffles.
  • Dynamic Partition Pruning (DPP): Skipping non-matching partition scans on Iceberg/Delta tables to minimize I/O overhead.
  • Shuffle & Memory Tuning: Reconfigured executor memory overhead, spark.sql.shuffle.partitions, and off-heap storage.

Workload Migration & Resilience Engineering

  • Kubernetes + MinIO Migration: Decoupled compute from storage, replacing monolithic on-prem YARN/HDFS clusters.
  • Dependency Isolation: Containerized Spark driver and executor environments preventing library clashes across jobs.
  • Automated Failure Recovery: Built DAG-level retry mechanisms and circuit breakers for transient network/disk failures.
  • Operational Runbooks: Standardized executor log debugging guides, cutting incident resolution time by 30%.
Tata Consultancy Services // Client Engagement

Nissan: Event-Driven Cloud Capture & Data Governance

Engineered serverless cloud ingestion leveraging AWS Lambda, Step Functions, and EventBridge handling 100K+ events daily with idempotent delivery and schema validation.

100K+
Daily Events Captured
High-frequency automotive telemetry and transactional events captured with zero message loss and idempotent processing.
99.5%
Output Reliability
Established programmatic data quality validation rules and version-controlled schema evolution across 20+ source systems.
Serverless Event-Driven Architecture (Nissan)
AWS Cloud Native
20+ SOURCES
Event Ingestion
Automotive systems emit events into Amazon EventBridge & SNS topic queues.
ORCHESTRATION
Step Functions & Lambda
Serverless state machines trigger idempotent Lambda functions for payload extraction.
GOVERNANCE
Schema Validation
Programmatic schema evolution handlers and quality check gates ensure 99.5% clean data.
MONITORING
Streamlit Dashboard
Interactive validation dashboard for operational monitoring and downstream consumer delivery.
Production Systems // Big Data & GenAI

End-to-End Lakehouses & Vector Retrieval Engines

Applying Lakehouse principles, SCD Type 2 dimensional modeling, and low-latency vector retrieval to real-world AI platforms.

Healthcare Lakehouse & RAG

AI Healthcare System

Python PySpark Databricks Delta Lake Airflow PostgreSQL Docker RAG
  • 500K+ Healthcare Records: Designed ETL pipelines using Delta Lake Medallion Architecture and SCD Type 2 modeling, optimizing partitioned storage and MERGE operations to reduce runtime by 40%.
  • FastAPI Serving: Delivered containerized FastAPI endpoints and PostgreSQL indexing to lower analytical query latency by 35%, scheduled via Airflow DAGs.
  • Sub-100ms Vector RAG: Implemented RAG pipeline integrating Cloudflare AI embeddings with a local vector cache, compressing context payload size by 80% and achieving sub-100ms response times.
GitHub
Streaming Engine & Vector Search

AI Recommendation Engine

Python PySpark Databricks Structured Streaming Delta Lake Airflow FastAPI
  • 20M+ Records Ingestion: Engineered batch and real-time ingestion flows using PySpark, SQL, and Structured Streaming on Databricks with Delta Lake checkpointing and incremental loading.
  • Sub-50ms Candidate Matching: Built a low-latency serving API using FastAPI and vector retrieval engine with pgvector/FAISS indexing across 100K+ items.
  • Automated Operations: Enforced data quality checks and automated scheduling via Apache Airflow and GitHub Actions CI/CD.
GitHub
Technical Matrix & Academic Honors

Technical Arsenal & Academic Foundation

Comprehensive technical skill matrix directly grounded in enterprise engineering and academic distinction.

Languages
Python SQL Scala Java
Big Data & Lakehouse
Apache Spark PySpark Spark SQL Databricks Unity Catalog Structured Streaming Apache Iceberg Delta Lake Medallion Arch MinIO / HDFS
Modeling & Quality
Dimensional Modeling SCD Type 2 Star Schema CDC Pipelines Schema Evolution Data Quality Rules
Cloud & Tooling
AWS (S3, Lambda) Step Functions EventBridge / SNS Apache Airflow AutoSys Docker GitHub Actions PostgreSQL FastAPI / Streamlit
Bachelor of Technology (2019 – 2023)

Guru Nanak Institutions Technical Campus

Computer Science and Engineering • Hyderabad, India

Major CGPA: 8.2 / 10.0 (First Class with Distinction) Minor in AI/ML: 8.6 / 10.0
Prestigious Award
Prime Minister's Scholarship Scheme (PMSS)
Awarded for academic excellence throughout B.Tech
Technical Discussion // Q&A Launchpad

Ready for Technical Drilldown

Click any topic below to pull up detailed architectural blueprints, trade-offs, and code-level decision rationale.

Spark Tuning & AQE

How broadcast joins, dynamic partition pruning, and shuffle partitions cut SLA runtime by 30%.

Open Architecture Blueprint

YARN to K8s + MinIO

Decoupling compute and storage, resolving dependency isolation, and saving 25% infra spend.

Open Migration Details

AWS Serverless Ingestion

100K+ daily events via Lambda, Step Functions, and EventBridge with idempotent delivery.

Open Cloud Blueprint

Medallion & SCD Type 2

Delta/Iceberg Bronze-Silver-Gold design, CDC handlers, and MERGE optimization.

Open Lakehouse Blueprint

GenAI & Vector Caching

Sub-100ms RAG retrieval using Cloudflare AI embeddings, FAISS indexing, and payload compression.

Open Vector Blueprint

Resilience & 30% MTTR Cut

Automated retries, fallback recovery handlers, executor log diagnostics, and operational runbooks.

Open Reliability Blueprint
TALKING POINTS
Introduce yourself: 3 years of experience as a Data Engineer at TCS building distributed Lakehouses & cloud pipelines across Nomura Capital and Nissan.