Eagleville, PA · Open to opportunities

Hey, I'm Phani K

I work with data for a living — cleaning it, questioning it, and turning it into answers people actually use. Most of my time goes into SQL, Python, pipelines, and dashboards that don't fall over.

profile.sql
1 row returned in 0.02s — scroll down to view full result set
← he follows your cursor, try it

// About

The short version.

I started out in India, picked up a Bachelor's in Computer Science, then moved to France for a Master's in Computer Engineering. These days I'm settled in Pennsylvania.

Somewhere along the way I figured out I enjoy the unglamorous side of data work — validating numbers before anyone sees them, reconciling datasets that refuse to match, and finding the one bad join that's been quietly wrecking a report for months. That instinct has carried me through healthcare claims, financial reporting, and a fair amount of pipeline plumbing.

Off the clock I tinker with side projects — most recently an iOS app, because apparently I don't stare at screens enough.

🇮🇳

India

🇫🇷

France

🇺🇸

USA

🎓 Master's in Computer EngineeringFrance
🎓 Bachelor's in Computer ScienceIndia
RIDGE PIKE EGYPT RD SCHUYLKILL RIVER
+
🇺🇸 Eagleville, PAPhani is here · open to remote

// What I do

Three hats, one job: make data useful.

📊

Data Analysis

Digging through messy datasets until they say something worth acting on.

  • Claims & fraud analytics
  • SQL ETL & anomaly detection
  • Forecasting & segmentation
⚙️

Data Engineering

Building the pipes so the analysis doesn't depend on someone's laptop.

  • ETL/ELT with Airflow & ADF
  • Streaming with Kafka & PySpark
  • Delta Lake & lakehouse patterns
📈

Business Intelligence

Dashboards people actually open — not the ones that die after the demo.

  • Power BI, Tableau, Looker Studio
  • KPI design & exec reporting
  • Excel automation (yes, still VBA)

// Toolbox

My stack, dropped on the floor.

Grab a tool and throw it around — they're sturdy, I use them daily.

physics on · drag anything

// Case studies

Work that shipped.

A mix of professional work (details anonymized) and portfolio builds. Filter by area.

Healthcare

Claims Analytics Dashboard

500K+ claims came in messy — duplicates, missing codes, odd billing patterns. Built SQL ETL to clean them, Power BI dashboards to track them, and Python checks to flag the weird ones. The anomaly flags ended up catching real revenue leakage.

500K+Claims cleaned
~$50KLeakage caught
12%Faster cycles
SQLPythonPower BIETL
Healthcare

Medicaid Claims Reporting

Multi-million-row Medicaid claims, strict compliance rules, and stakeholders who lived in Excel. Built SQL Server data models for clinical and financial reporting, plus VBA-powered self-serve tools so people stopped emailing me for numbers.

60%Less manual work
HIPAAAware handling
SQL ServerT-SQLPythonExcel/VBA
Finance

Fraud Detection Analysis

The fraud queue was drowning reviewers in false positives. Went through 100K+ transactions, found the payment patterns that mattered, and built Python models plus Tableau monitoring that let reviewers focus on the cases actually worth their time.

18%Fewer false positives
100K+Transactions
SQLPythonTableau
Finance

Financial Reporting Automation

Monthly reporting used to be a week of copy-paste. Replaced it with Python and SQL pipelines feeding Power BI dashboards for KPIs and portfolio risk. The stored procedures are boring — which is exactly what you want from month-end.

15+ hrsSaved per cycle
PythonT-SQLPower BIExcel
Finance

Real-Time Operations Monitoring

Live data feeds where a bad number published downstream means a bad day. Built validation checks that flag out-of-range values before publication, with dashboards and alerting tuned to catch issues inside tight SLA windows.

Pre-publishValidation
SLAWindows held
PythonSQLPower BIAlerting
Product

Campaign ROI Analysis

Ad budget was being spread evenly across audiences that performed very differently. Segmented 100K+ customer interactions, compared ROI per segment, and moved spend where it worked. Simple idea, real savings.

25%More conversions
15%Budget saved
SQLPythonPower BIGA4
Product

Churn & Adoption Forecasting

Built SQL datasets (heavy on window functions) and Python/R models over 1M+ records to predict churn and device adoption. The forecasts got noticeably better than gut feel, and the churn flags surfaced customer segments worth saving.

25%Forecast accuracy gain
1M+Records modeled
SQLPythonRLooker Studio
Product

A/B Testing & Funnel Analysis

Instrumented an onboarding funnel, tracked where users dropped, and ran controlled experiments on the worst steps. Learned the hard way that most A/B tests are inconclusive — the discipline is in the readout, not the test.

FunnelDrop-off mapped
A/BExperiment readouts
SQLPythonEvent dataStatistics
Engineering

Healthcare ETL Pipeline

End-to-end pipeline simulating Medicaid claims processing: Python extract/transform, PostgreSQL storage, Airflow scheduling, and HIPAA-aware handling baked into the design. Handles 1M+ records a month without drama.

1M+Monthly records
AutoValidation & scheduling
PythonPostgreSQLAirflowDocker
↗ github.com/phanikan
Engineering

Fraud Detection Streaming Pipeline

Real-time take on the fraud problem: Kafka ingestion, PySpark feature engineering, Delta Lake storage, and statistical thresholds that flag suspicious transactions as they happen — with monitoring so I know when it breaks before anyone else does.

Real-timeStreaming
LiveAlerts
KafkaPySparkDelta LakeSpark Streaming
↗ github.com/phanikan
Engineering

Azure Lakehouse Pipeline

Bronze/silver/gold medallion build on Azure Data Factory and Databricks, moving 2M+ records a day through automated quality checks with CI/CD deployment. The 40% runtime cut came from fixing partitioning, not from adding hardware.

2M+Daily records
40%Faster processing
Azure ADFDatabricksPySparkDelta LakeGit
Engineering

Master Data Management & Migration

500K+ device and CRM records that needed to survive a stage-to-production migration intact. Deduplication, validation rules, integrity checks — the unglamorous work that decides whether a migration is a milestone or an incident.

500K+Records migrated
99%+Clean readiness
SQLExcelData governanceMigration
Side project

Aervoya — iOS Travel App

A Swift/SwiftUI app with live flight tracking off the AviationStack API. Nothing to do with my day job, which is exactly why I built it — sometimes you learn the most from the stack you don't use at work.

SwiftSwiftUIREST APIsiOS
↗ github.com/phanikan

// Impact

Proof over promises.

records / day2M+through the Azure lakehouse pipeline, bronze to gold
leakage caught~$50Krevenue leakage flagged by claims anomaly checks
manual work cut60%drop in manual claims processing after pipeline automation
records / month1M+through the healthcare ETL pipeline on Airflow
false positives−18%fewer bad flags reaching fraud reviewers
per cycle15+ hrssaved every reporting cycle after automating month-end
claims cleaned500K+standardized in one SQL ETL build
etl runtime−40%faster runs — fixed partitioning, not more hardware
forecasts+25%accuracy gain from churn & adoption models
dataset prephrs→15mafter prototyping the migration pipeline
dataset readiness99%+after dedup and validation across 500K+ records
transactions100K+analyzed for fraud payment patterns
ad budget15%recovered by moving spend to segments that converted

Hover a number — every one of them was delivered, not projected.

// Contact

Got a data problem worth solving?

I'd like to hear about it. Email works best — I actually read it.