Saikrishna Poluri
Lead Data Engineer — Azure · Microsoft Fabric · Databricks · Power BI · GenAI. 9+ years architecting enterprise data platforms; open to Lead and Staff Data Engineer, Azure / Microsoft Fabric, and Databricks roles.
I design, build, and run production data platforms across the Microsoft Azure ecosystem — Azure Data Factory, ADLS Gen2, Azure Synapse, Microsoft Fabric, and Azure Databricks. I build Bronze/Silver/Gold medallion pipelines on Delta Lake with governance and lineage, design Power BI semantic models for executive reporting, and back every release with CI/CD (Azure DevOps, GitHub), monitoring, runbooks, and FinOps discipline. In the past year I also shipped three production GenAI agents on the lakehouse. My domain depth is enterprise finance — Record-to-Report, P&L, Plan/Budget/Forecast/Actuals — across SAP, JDE, Hyperion HFM, IBM DB2, and PeopleSoft, with a track record of explaining architecture tradeoffs to CFO-level and executive stakeholders.
sk.poluri8@gmail.com · LinkedIn · GitHub · TX, US · open to relocation
Flagship work
ADM — Finance R2R Lakehouse & Production GenAI Agents
What: I lead end-to-end delivery of Finance Record-to-Report analytics at Archer Daniels Midland. I architected medallion pipelines on Azure Databricks and ADF that integrate JDE, SAP, HFM, IBM DB2, and IBM APGO Web Wire into Delta Lake with lineage and governance, feeding Power BI semantic models and KPI frameworks (P&L, Budget vs. Actual, Expense Forecasting, EHS, Capital Allocation). On the same platform I built and shipped three production GenAI agents over PeopleSoft and SAP labor data: a Genie Space data agent, a Vector Search RAG assistant, and a multi-agent router — hosted as a Databricks App with row- and column-level security.
Impact: Weekly-to-yearly reporting for the Executive Committee and CFO-level consumers; the GenAI agents serve 100+ users (including ED/VP/C-suite) over a 10,000+ employee global workforce dataset, handling 1,000+ queries per day across 30+ policy documents. Drove an estimated ~$500K in annual compute savings through FinOps tuning.
Maersk — Fact-Based Reporting & Dremio Semantic Layer
What: I delivered end-to-end data engineering for Maersk's Fact-Based Reporting program: metadata-driven ADF and PySpark ingestion of SAP S/4HANA, SAP AC-DOCA Universal Journal, and Hyperion HFM into Delta on ADLS Gen2, modeling in Databricks, publication to Azure Synapse Dedicated SQL Pool, and migration of SSAS multidimensional cubes to Azure Analysis Services Tabular. I then implemented a Dremio and Power BI semantic-layer lakehouse and moved managed tables to external tables for governance and cost control.
Impact: A multi-million-dollar annual infrastructure cost reduction from the semantic-layer lakehouse, faster CFO-level dashboards, and lower-latency Finance analyst exploration on Actuals, Plan, and Forecast P&L.
PANYNJ — ICMS Financial Lakehouse & PATH Ridership Migration
What: For the Port Authority of NY & NJ I designed metadata-driven ADF and PySpark ingestion into ADLS Gen2 for the agency-wide Integrated Cost Management System, integrating SAP ECC FI, IBM Planning Analytics, and Budget PRO for Plan, Budget, Forecast, and Actuals reporting across Capital and Operating portfolios. In parallel I migrated legacy VBA and MS Access PATH ridership workloads to parameterized Azure Synapse pipelines with Power BI dashboards.
Impact: Agency-wide and division-level financial portfolio visibility for Finance leadership and retirement of fragile desktop tooling — all within a five-month contract.
ADM Treasury — “Forecasting the Flow” (Hackathon: 2nd place, Bard’s Tale Award)
What: On an 11-person hackathon team I built the lakehouse and data/feature pipeline behind a probabilistic 90-day cash-flow forecasting platform for ADM Treasury, and served as Lead Presenter on Demo Day. I ingested 15 Agris ERP tables plus weather, corn-futures, and USDA feeds into a Bronze/Silver/Gold medallion, and engineered the governed feature views that fed the models. I contributed to the model design alongside our ML engineers, who built the cascaded quantile-regression models (XGBoost/LightGBM).
Impact: The team placed 2nd of the field and won the Bard’s Tale Award; delivery-timing accuracy 26–36% better than a naive baseline, with ~80% of outcomes inside the forecast interval — giving Treasury planning ranges instead of a single, brittle point estimate.
Experience
Data & Analytics Technical Lead
- Lead technical planning and end-to-end delivery for enterprise Finance Record-to-Report analytics: ingestion, lakehouse transformation, semantic modeling, reporting, stakeholder alignment, runbooks, monitoring, and hypercare.
- Architected Bronze/Silver/Gold pipelines on Azure Databricks and ADF integrating JDE, SAP, HFM, IBM DB2, and IBM APGO Web Wire into Delta Lake with lineage and governance; delivered Power BI semantic models and KPI frameworks for Executive Committee and CFO-level consumers.
- Built and deployed three production GenAI agents (Genie Space data agent, Vector Search RAG assistant, multi-agent router) serving 100+ users over a 10,000+ employee dataset with row- and column-level security.
- Drove cost and performance optimization (cluster right-sizing, autoscaling, Delta tuning) for an estimated ~$500K in annual compute savings; automated orchestration across ADF, Databricks Jobs, and Power BI refreshes with CI/CD gates and monitoring in Azure DevOps.
- Provide fit-gap and architecture recommendations and mentor junior engineers through design reviews, code reviews, and PySpark best practices.
Senior Azure Data Engineer (Contract)
- Designed metadata-driven ADF and PySpark ingestion frameworks feeding ADLS Gen2 for the agency-wide Integrated Cost Management System financial lakehouse.
- Integrated SAP ECC FI, IBM Planning Analytics, and Budget PRO for Plan, Budget, Forecast, and Actuals reporting across Capital and Operating portfolios.
- Migrated legacy VBA and MS Access PATH ridership workloads to parameterized Azure Synapse pipelines and Power BI.
Azure Data Engineer, Fact-Based Reporting
- Delivered end-to-end data engineering for the Fact-Based Reporting program, integrating SAP S/4HANA, SAP AC-DOCA Universal Journal, and Hyperion HFM into an Azure lakehouse for Actuals, Plan, and Forecast P&L reporting.
- Built metadata-driven ADF and PySpark ingestion to Delta on ADLS Gen2 and published refined datasets to Azure Synapse Dedicated SQL Pool; migrated SSAS multidimensional cubes to Azure Analysis Services Tabular.
- Implemented a Dremio and Power BI semantic-layer lakehouse credited with a multi-million-dollar annual infrastructure cost reduction.
Senior Data Engineer, Customer Data Management
- Built Customer Data Management solutions on Azure using ADF, SSIS, PySpark, and Azure Databricks.
- Customized TIBCO EBX for Master Data Management and executed the EBX 5.9 to 6.0.11 upgrade with zero downstream disruption.
- Built a Power BI dashboard with embedded Power Apps for golden-copy/survivorship stewardship; optimized Azure infrastructure for recurring monthly cost savings.
Senior Analytics Consultant, Worldwide Payment Services
- Migrated on-premises SQL Server databases and SSIS workloads to Azure SQL Database and ADF using the SSIS Integration Runtime.
- Rebuilt legacy SSRS reporting in Power BI; implemented U-SQL transformations in Azure Data Lake Analytics; developed dimensional models and SCD 0/1/2 patterns.
Software Engineer, Data Warehousing & ETL
- Designed and deployed SSIS packages, SSRS reports, and T-SQL stored procedures, views, and triggers for lottery data warehousing.
- Managed SQL Server Agent scheduling, audit logging, and data validation; performed root-cause analysis and performance tuning for production ETL.
Certifications
Microsoft Certified: Fabric Data Engineer Associate
DP-700 — earned.
Microsoft Certified: Fabric Analytics Engineer Associate
DP-600 — earned.
Microsoft Certified: Azure Data Engineer Associate
DP-203 — earned.
Microsoft Certified: Azure Data Fundamentals / Fundamentals
DP-900 & AZ-900 — earned.
Databricks Lakehouse Fundamentals · Dremio Verified Lakehouse Associate
Earned.
Databricks Certified Data Engineer (Associate & Professional) · GenAI Engineer Associate
In progress — targeting 2026.
Skills
| Area | Skills |
|---|---|
| Azure & Microsoft Fabric | Azure Data Factory, ADLS Gen2, Azure Synapse Analytics (Notebooks, SQL Analytics endpoints, dedicated & serverless pools), Azure SQL, Azure Analysis Services, Microsoft Fabric (Lakehouse, Warehouse, Pipelines, Dataflows, Semantic Models, DirectLake), Azure DevOps |
| Data Engineering & Lakehouse | Delta Lake, Bronze/Silver/Gold medallion, dimensional & semantic modeling, CDC, SCD 0/1/2, schema evolution, metadata-driven ingestion, data quality, managed vs. external tables, lineage and governance |
| Languages | SQL, T-SQL, Spark SQL, Python, PySpark, U-SQL |
| BI & Semantic Layers | Power BI, DAX, semantic models, DirectLake & report performance tuning, SSAS Tabular and Multidimensional, Dremio, Denodo, Power Apps, Power Automate |
| Security, CI/CD & DevOps | RBAC, Azure Key Vault, network security (VNet / Private Link), Git, Azure DevOps CI/CD, GitHub Actions, Terraform (working knowledge), monitoring, runbooks, hypercare |
| Databricks & GenAI | Azure Databricks, Apache Spark, Unity Catalog, Lakeflow Jobs, Mosaic AI Vector Search, AI/BI Genie, Databricks Apps, RAG, multi-agent orchestration, MLflow, Anthropic Claude API, MCP |
| ERP & Enterprise Systems | SAP S/4HANA, SAP ECC FI, SAP BW, SAP AC-DOCA Universal Journal, JDE, Hyperion HFM, IBM DB2, IBM Planning Analytics, PeopleSoft, TIBCO EBX, Budget PRO |
What I bring
Engineering depth across Azure & Databricks
- Production Spark and PySpark at enterprise scale, with Delta Lake design covering CDC, SCD 0/1/2, schema evolution, compaction, and performance tuning.
- Broad Microsoft data platform: Azure Data Factory, Synapse, Microsoft Fabric, ADLS Gen2, Azure SQL, and Analysis Services, plus Power BI and external semantic layers (Dremio/Denodo).
- Real orchestration and DevOps: Databricks Jobs coordinated with ADF and Redwood RunMyJobs, gated by Azure DevOps CI/CD with monitoring, runbooks, and hypercare.
- FinOps evidence: cluster right-sizing and autoscaling at ADM (~$500K annual savings), the Maersk managed-to-external table shift, and recurring Azure cost savings at Ecolab.
Delivery, architecture & leadership
- End-to-end delivery ownership: technical planning, cross-team dependencies, stakeholder alignment, and hypercare for CFO-level Finance analytics.
- Architecture advisory in practice: fit-gap recommendations on batch vs. near-real-time, CDC/SCD design, and semantic-layer selection.
- Governance over sensitive data: lineage plus row- and column-level controls protecting HR and finance datasets in regulated environments.
- Production GenAI: Genie, Vector Search RAG, multi-agent routing, and Databricks Apps — shipped, governed, and in daily use.
- Mentoring and enablement: design reviews, code reviews, and engineering standards for junior engineers.
Education
Bachelor of Technology (B.Tech), Electronics & Communication Engineering — Jawaharlal Nehru Technological University, Anantapuramu, India, 2015.