Hey, I'm Phani K
I work with data for a living — cleaning it, questioning it, and turning it into answers people actually use. Most of my time goes into SQL, Python, pipelines, and dashboards that don't fall over.
// About
The short version.
I started out in India, picked up a Bachelor's in Computer Science, then moved to France for a Master's in Computer Engineering. These days I'm settled in Pennsylvania.
Somewhere along the way I figured out I enjoy the unglamorous side of data work — validating numbers before anyone sees them, reconciling datasets that refuse to match, and finding the one bad join that's been quietly wrecking a report for months. That instinct has carried me through healthcare claims, financial reporting, and a fair amount of pipeline plumbing.
Off the clock I tinker with side projects — most recently an iOS app, because apparently I don't stare at screens enough.
India
France
USA
// What I do
Three hats, one job: make data useful.
Data Analysis
Digging through messy datasets until they say something worth acting on.
- Claims & fraud analytics
- SQL ETL & anomaly detection
- Forecasting & segmentation
Data Engineering
Building the pipes so the analysis doesn't depend on someone's laptop.
- ETL/ELT with Airflow & ADF
- Streaming with Kafka & PySpark
- Delta Lake & lakehouse patterns
Business Intelligence
Dashboards people actually open — not the ones that die after the demo.
- Power BI, Tableau, Looker Studio
- KPI design & exec reporting
- Excel automation (yes, still VBA)
// Toolbox
My stack, dropped on the floor.
Grab a tool and throw it around — they're sturdy, I use them daily.
// Case studies
Work that shipped.
A mix of professional work (details anonymized) and portfolio builds. Filter by area.
Claims Analytics Dashboard
500K+ claims came in messy — duplicates, missing codes, odd billing patterns. Built SQL ETL to clean them, Power BI dashboards to track them, and Python checks to flag the weird ones. The anomaly flags ended up catching real revenue leakage.
Medicaid Claims Reporting
Multi-million-row Medicaid claims, strict compliance rules, and stakeholders who lived in Excel. Built SQL Server data models for clinical and financial reporting, plus VBA-powered self-serve tools so people stopped emailing me for numbers.
Fraud Detection Analysis
The fraud queue was drowning reviewers in false positives. Went through 100K+ transactions, found the payment patterns that mattered, and built Python models plus Tableau monitoring that let reviewers focus on the cases actually worth their time.
Financial Reporting Automation
Monthly reporting used to be a week of copy-paste. Replaced it with Python and SQL pipelines feeding Power BI dashboards for KPIs and portfolio risk. The stored procedures are boring — which is exactly what you want from month-end.
Real-Time Operations Monitoring
Live data feeds where a bad number published downstream means a bad day. Built validation checks that flag out-of-range values before publication, with dashboards and alerting tuned to catch issues inside tight SLA windows.
Campaign ROI Analysis
Ad budget was being spread evenly across audiences that performed very differently. Segmented 100K+ customer interactions, compared ROI per segment, and moved spend where it worked. Simple idea, real savings.
Churn & Adoption Forecasting
Built SQL datasets (heavy on window functions) and Python/R models over 1M+ records to predict churn and device adoption. The forecasts got noticeably better than gut feel, and the churn flags surfaced customer segments worth saving.
A/B Testing & Funnel Analysis
Instrumented an onboarding funnel, tracked where users dropped, and ran controlled experiments on the worst steps. Learned the hard way that most A/B tests are inconclusive — the discipline is in the readout, not the test.
Healthcare ETL Pipeline
End-to-end pipeline simulating Medicaid claims processing: Python extract/transform, PostgreSQL storage, Airflow scheduling, and HIPAA-aware handling baked into the design. Handles 1M+ records a month without drama.
Fraud Detection Streaming Pipeline
Real-time take on the fraud problem: Kafka ingestion, PySpark feature engineering, Delta Lake storage, and statistical thresholds that flag suspicious transactions as they happen — with monitoring so I know when it breaks before anyone else does.
Azure Lakehouse Pipeline
Bronze/silver/gold medallion build on Azure Data Factory and Databricks, moving 2M+ records a day through automated quality checks with CI/CD deployment. The 40% runtime cut came from fixing partitioning, not from adding hardware.
Master Data Management & Migration
500K+ device and CRM records that needed to survive a stage-to-production migration intact. Deduplication, validation rules, integrity checks — the unglamorous work that decides whether a migration is a milestone or an incident.
Aervoya — iOS Travel App
A Swift/SwiftUI app with live flight tracking off the AviationStack API. Nothing to do with my day job, which is exactly why I built it — sometimes you learn the most from the stack you don't use at work.
↗ github.com/phanikan// Impact
Proof over promises.
Hover a number — every one of them was delivered, not projected.
// Contact
Got a data problem worth solving?
I'd like to hear about it. Email works best — I actually read it.