About Experience Projects Skills Contact
Data Engineer · Bengaluru, India

ADITH
KRISHNA

Building Azure data pipelines that scale,
with precision, reliability, and zero drama.

2+ years turning messy, high-volume data into clean analytics infrastructure. Specialising in ADF, Microsoft Fabric, Synapse, PySpark, and the full Azure modern data stack — from raw ingestion to Power BI dashboards.

Live Pipeline Topology
Raw Sources
ADF
ADLS Gen2
PySpark
Synapse
Power BI
● Medallion Architecture — Bronze → Silver → Gold
Impact Snapshot
30%
Spark job speedup via query optimisation
50%
Query latency drop on Gold layer
60%
Fewer data-quality incidents downstream
~45m
Batch time (was 6hrs) after migration
0yr
Experience
0+
Pipelines Delivered
0×
Azure Services
0+
Projects Shipped
End-to-End Data Architecture Stack
📥
Raw Sources
Ingestion
🏭
ADF Pipelines
Orchestration
🏔️
ADLS Gen2
Bronze Layer
PySpark / Fabric
Silver Layer
🗄️
Synapse SQL
Gold Layer
🔍
Data Quality
Observability
📊
Power BI
Reporting
00 / Who

About Me

I'm a Data Engineer at LTIMindtree, Bengaluru, where I design and build cloud-native data infrastructure on the Azure ecosystem — primarily for the Microsoft Learn platform analytics workload.

My day-to-day involves translating complex, high-volume raw data flows into clean, reliable pipelines that business stakeholders can actually trust. That means ADF orchestration, PySpark transformations, Medallion Architecture on Microsoft Fabric, and Synapse Analytics SQL pools.

I care deeply about pipeline resilience — incremental loads, watermarking, trigger-based scheduling, SLA monitoring, and automated data-quality checks baked in from the start, not bolted on at the end.

Before that, I trained through the Great Learning Academy Software Development Program and completed a B.Tech in Computer Science at VIT Chennai. I hold a Microsoft AZ-900 certification and am continuously expanding deeper into the Azure data engineering certifications track.

Outside work: I enjoy reading about distributed systems, experimenting with Delta Lake, and contributing to internal knowledge-sharing sessions at my team.

B.Tech Computer Science
Vellore Institute of Technology, Chennai
2019 – 2023
Software Development Program
Great Learning Academy
Sep 2023 – Jun 2024
🏆
Microsoft Certified: Azure Fundamentals
AZ-900 · Microsoft
01 / Work

Experience

LTIMindtree
May 2024 — Present
📍 Bengaluru, India
DATA
ENGINEER
Microsoft Learn Platform Analytics
Microsoft Skilling Track
  • Designed and maintained end-to-end ETL pipelines using ADF, Microsoft Fabric, and Synapse Analytics to ingest and process large-scale learning activity data — millions of events daily from the Microsoft Learn platform.
  • Built PySpark transformation workflows for structured and semi-structured datasets; optimised Spark jobs and SQL queries, cutting processing time by ~30% through partition pruning, broadcast joins, and predicate pushdown.
  • Implemented incremental loads, trigger-based scheduling, and validation checkpoints; monitored pipeline health with Azure Monitor alerts and resolved failures proactively to meet SLA targets.
  • Collaborated directly with Microsoft client teams to define schema contracts, data models, and reporting requirements — translating business requirements into reliable data products.
Enterprise Data Integration
Internal Platform
  • Developed scalable ingestion pipelines integrating multiple internal data sources into ADLS Gen2 using reusable ADF templates — reducing new-source onboarding effort by ~40% per source.
  • Enforced data quality rules and schema standardisation across Parquet, CSV, and JSON datasets with PySpark-based validation hooks, ensuring analytics layer reliability for downstream consumers.
  • Built CI/CD pipelines in Azure DevOps for ADF deployments across Dev, UAT, and Prod environments — enabling parameterised, safe releases with automated testing and rollback gates.
02 / Build

Projects

01
🏗️
Real-Time Learning Analytics Platform
Full Medallion Architecture on Microsoft Fabric Lakehouses — Bronze ingestion to Gold reporting — with ADF event-trigger pipelines, PySpark Silver-layer cleansing, and Synapse dedicated SQL pool for Gold reporting tables powering live Power BI dashboards.
  • Cut reporting query latency by ~50% with partitioned and clustered Gold-layer tables
  • Power BI dashboards with row-level security served to 5+ stakeholder teams across regions
  • Event-trigger ingestion handles clickstream bursts with zero manual intervention
ADFMicrosoft Fabric PySparkSynapse SQL Power BIADLS Gen2 Delta Lake
02
🔍
Data Quality & Observability Framework
Metadata-driven PySpark quality framework covering schema validation, null checks, referential integrity, and range assertions — all configured via JSON rule files in ADLS, integrated as modular ADF pipeline activities with automated Azure Monitor alerting and a Synapse audit table.
  • Reduced data-quality incidents escalated downstream by ~60%
  • Zero-code rule updates — new validation rules via JSON config with no pipeline redeployment
  • Full audit trail in Synapse SQL with every failure logged, timestamped, and alerted
PySparkADF Azure MonitorSynapse SQL PythonJSON Config
03
☁️
Incremental ETL Migration — On-Prem to Azure
Migrated 15+ legacy SQL Server batch jobs to cloud-native ADF incremental pipelines using watermark and change-tracking patterns, with full CI/CD across Dev/UAT/Prod in Azure DevOps. Eliminated overnight batch windows and dramatically reduced infrastructure costs.
  • Batch processing time slashed from ~6 hours to ~45 minutes
  • Watermark & change-tracking incremental patterns for 15+ source tables
  • Azure DevOps CI/CD with parameter overrides, gate approvals per environment
ADFADLS Gen2 Azure SQLAzure DevOps ParquetCI/CD
03 / Stack

Skills

☁️
Cloud — Azure
Data Factory Microsoft Fabric Synapse Analytics ADLS Gen2 Azure Monitor Azure DevOps Azure SQL DB
Languages & Big Data
PySpark Python SQL Apache Spark Spark Optimisation Partition Tuning
🗄️
Data Engineering
ETL / ELT Medallion Architecture Incremental Loading Data Warehousing Dimensional Modelling Data Quality Schema Evolution
📦
Formats & Databases
Parquet Delta Lake JSON / CSV SQL Server Synapse SQL Pool Azure SQL DB
🔧Git
🚀Azure DevOps
🔄CI/CD
🐧Linux
💻VS Code
📓Jupyter
🏆AZ-900 Certified

LET'S
BUILD
TOGETHER.

Open to data engineering roles. If you're working on interesting data problems at scale — pipelines, lakehouses, real-time or batch — let's talk.