Saikrishna Poluri
Lead Data Engineer / Resident Solutions Architect — Databricks, Lakehouse & GenAI. 9+ years architecting enterprise data platforms, now building production lakehouses and GenAI agents on the Databricks Data Intelligence Platform.
I design, build, and run production lakehouses on the Databricks Data Intelligence Platform — Bronze/Silver/Gold medallion pipelines on Delta Lake with Unity Catalog governance and lineage, orchestrated across Databricks Jobs, Azure Data Factory, and Redwood RunMyJobs, with CI/CD, monitoring, runbooks, and FinOps discipline behind every release. In the past year I shipped three production GenAI agents on Databricks — a Genie Space data agent, a Vector Search RAG assistant, and a multi-agent router, delivered as a governed Databricks App with Unity Catalog row- and column-level security. My domain depth is enterprise finance — Record-to-Report, P&L, Plan/Budget/Forecast/Actuals — across SAP, JDE, Hyperion HFM, IBM DB2, and PeopleSoft, with a track record of explaining architecture tradeoffs to CFO-level stakeholders. I'm targeting Databricks Data Engineer and Resident Solutions Architect roles, and I'm open to any location.
sk.poluri8@gmail.com · LinkedIn · GitHub · TX, US
Flagship work
ADM — Finance R2R Lakehouse & Production GenAI Agents
What: I lead end-to-end delivery of Finance Record-to-Report analytics at Archer Daniels Midland. I architected medallion pipelines on Azure Databricks and ADF that integrate JDE, SAP, HFM, IBM DB2, and IBM APGO Web Wire into Delta Lake with lineage and Unity Catalog governance, feeding KPI frameworks for P&L, Budget vs. Actual, Expense Forecasting, EHS, and Capital Allocation. On the same platform I built and shipped three production GenAI agents over PeopleSoft and SAP labor data: a Corporate Headcount & Salaries Data Agent in a Genie Space (SQL shown with every answer for trust), an Employee Policy RAG Assistant on Vector Search and LLM endpoints with cited answers, and a Combined Multi-Agent Assistant with intent routing — hosted as a Databricks App with a custom Streamlit UI, embedded Genie Space, Databricks AI/BI dashboards, and Unity Catalog row- and column-level permissions for HR users.
Impact: Weekly-to-yearly reporting for the Executive Committee and CFO-level consumers, plus governed natural-language self-service over sensitive HR data — all in production today.
Maersk — Fact-Based Reporting & Dremio Semantic Layer
What: I delivered end-to-end data engineering for Maersk's Fact-Based Reporting program: metadata-driven ADF and PySpark ingestion of SAP S/4HANA, SAP AC-DOCA Universal Journal, and Hyperion HFM data into Delta on ADLS Gen2, financial dimension and Universal Journal modeling in Databricks, publication to Azure Synapse Dedicated SQL Pool, and migration of SSAS multidimensional cubes to Azure Analysis Services Tabular. I then implemented a Dremio and Power BI semantic-layer lakehouse and moved managed tables to external tables for governance and cost control.
Impact: A multi-million-dollar annual infrastructure cost reduction from the semantic-layer lakehouse, faster CFO-level dashboards, and lower-latency Finance analyst exploration on Actuals, Plan, and Forecast P&L.
PANYNJ — ICMS Financial Lakehouse & PATH Ridership Migration
What: For the Port Authority of NY & NJ I designed metadata-driven ADF and PySpark ingestion into ADLS Gen2 Bronze for the agency-wide Integrated Cost Management System, integrating SAP ECC FI, IBM Planning Analytics, and Budget PRO to support Plan, Budget, Forecast, and Actuals reporting across Capital and Operating portfolios for Aviation, PATH, and Tunnels, Bridges & Terminals. In parallel I migrated legacy VBA and MS Access PATH ridership workloads to parameterized Azure Synapse pipelines with Power BI dashboards.
Impact: Agency-wide and division-level financial portfolio visibility for Finance leadership and retirement of fragile desktop tooling — all within a five-month contract.
ADM Treasury — “Forecasting the Flow” (Hackathon: 2nd place, Bard’s Tale Award)
What: On an 11-person hackathon team I built the Databricks lakehouse and data/feature pipeline behind a probabilistic 90-day cash-flow forecasting platform for ADM Treasury, and served as Lead Presenter on Demo Day. I ingested 15 Agris ERP tables plus Maxar weather, corn-futures, and USDA feeds into a Bronze/Silver/Gold medallion under Unity Catalog, and engineered the governed feature and treasury-facing views that fed the forecasting models. I contributed to the model design alongside our ML engineers, who built the cascaded quantile-regression models (XGBoost/LightGBM) that turn those features into P10/P50/P90 cash-timing intervals.
Impact: The team placed 2nd of the field and won the Bard’s Tale Award; delivery-timing MAE of 6.5–7.6 days (26–36% better than a naive baseline), with ~80% of outcomes inside the P10–P90 band — giving Treasury planning intervals instead of a single, brittle point estimate.
Experience
Data & Analytics Technical Lead
- Lead technical planning and end-to-end delivery for enterprise Finance Record-to-Report analytics: ingestion, lakehouse transformation, semantic modeling, reporting, stakeholder alignment, runbooks, monitoring, and hypercare.
- Architected Bronze/Silver/Gold pipelines on Azure Databricks and ADF integrating JDE, SAP, HFM, IBM DB2, and IBM APGO Web Wire into Delta Lake with lineage and Unity Catalog governance; delivered P&L, Budget vs. Actual, Expense Forecasting, EHS, and Capital Allocation KPI frameworks for Executive Committee and CFO-level consumers.
- Built and deployed three production GenAI agents on Databricks (Genie Space data agent, Vector Search RAG assistant, multi-agent router) hosted as a Databricks App with Streamlit UI and Unity Catalog row- and column-level permissions.
- Drive Databricks FinOps and performance work — cluster right-sizing, autoscaling policies, and Delta tuning (partitioning, compaction, Z-ORDER) — and automate orchestration across ADF, Databricks Lakeflow Jobs, and Power BI refreshes via Redwood RunMyJobs, with CI/CD gates and monitoring in Azure DevOps.
- Provide fit-gap and architecture recommendations (batch vs. near-real-time, CDC/SCD design, Power BI vs. Dremio/Denodo semantic layers) and mentor junior engineers through design reviews, code reviews, and PySpark best practices.
Senior Azure Data Engineer (Contract)
- Designed metadata-driven ADF and PySpark ingestion frameworks feeding ADLS Gen2 Bronze for the agency-wide Integrated Cost Management System financial lakehouse.
- Integrated SAP ECC FI, IBM Planning Analytics, and Budget PRO for Plan, Budget, Forecast, and Actuals reporting across Capital and Operating portfolios spanning Aviation, PATH, and Tunnels, Bridges & Terminals.
- Migrated legacy VBA and MS Access PATH ridership workloads to parameterized Azure Synapse pipelines and Power BI.
Azure Data Engineer, Fact-Based Reporting
- Delivered end-to-end data engineering for the Fact-Based Reporting program, integrating SAP S/4HANA, SAP AC-DOCA Universal Journal, and Hyperion HFM into an Azure lakehouse for Actuals, Plan, and Forecast P&L reporting.
- Built metadata-driven ADF and PySpark ingestion to Delta on ADLS Gen2, modeled SAP financial dimensions and Universal Journal line items in Databricks, and published refined datasets to Azure Synapse Dedicated SQL Pool.
- Migrated SSAS multidimensional cubes to Azure Analysis Services Tabular, improving query performance for CFO-level dashboards and Finance analyst exploration.
- Implemented a Dremio and Power BI semantic-layer lakehouse credited with a significant, multi-million-dollar annual infrastructure cost reduction.
- Moved Databricks workloads from managed to external table architecture for stronger governance and cost management across Finance datasets.
Senior Data Engineer, Customer Data Management
- Built Customer Data Management solutions on Azure using ADF, SSIS, PySpark, and Azure Databricks.
- Customized TIBCO EBX for Master Data Management and executed the EBX 5.9 to 6.0.11 upgrade with zero downstream disruption.
- Optimized Azure infrastructure for meaningful recurring monthly cost savings.
- Built a Power BI dashboard with embedded Power Apps so data stewards could review and approve golden-copy/survivorship records for match-merge exceptions.
Senior Analytics Consultant, Worldwide Payment Services
- Migrated on-premises SQL Server databases and SSIS workloads to Azure SQL Database and ADF using the SSIS Integration Runtime.
- Rebuilt legacy SSRS reporting in Power BI with interactive, self-service dashboards; implemented U-SQL transformations in Azure Data Lake Analytics.
- Developed dimensional models and SCD 0/1/2 patterns for payment-services data warehouse reporting.
Software Engineer, Data Warehousing & ETL
- Designed and deployed SSIS packages, SSRS reports, and T-SQL stored procedures, views, and triggers for lottery data warehousing.
- Built and maintained data warehouse ETL for Maryland Lottery and Atlantic Lottery reporting.
Certifications
Microsoft Certified: Fabric Data Engineer Associate
DP-700 — earned.
Microsoft Certified: Fabric Analytics Engineer Associate
DP-600 — earned.
Microsoft Certified: Azure Data Engineer Associate
DP-203 — earned.
Microsoft Certified: Azure Data Fundamentals
DP-900 — earned.
Microsoft Certified: Azure Fundamentals
AZ-900 — earned.
Databricks Lakehouse Fundamentals
Databricks accreditation — earned.
Dremio Verified Lakehouse Associate
Dremio — earned.
Databricks Certified Data Engineer Associate
In progress — planned, target 2026. Not yet earned; this knowledge base is my structured preparation for it.
Skills
| Area | Skills |
|---|---|
| Databricks & Spark | Azure Databricks, Apache Spark, PySpark, Spark SQL, Lakeflow Jobs (formerly Databricks Workflows), AI/BI Genie, Databricks Apps, Mosaic AI Vector Search, MLflow, Mosaic AI Model Serving, Delta performance tuning, medallion architecture |
| Lakehouse & Governance | Delta Lake, Bronze/Silver/Gold design, Unity Catalog (lineage, row- and column-level security), managed vs. external tables, CDC, SCD 0/1/2, schema evolution, metadata-driven ingestion, data quality, dimensional and semantic modeling |
| Cloud & Orchestration | Microsoft Azure, Azure Data Factory, Azure Synapse Analytics, Microsoft Fabric, ADLS Gen2, Azure SQL Database, Azure Analysis Services, Redwood RunMyJobs, Azure DevOps CI/CD, GitHub Actions, Terraform (working knowledge), monitoring, runbooks, hypercare |
| GenAI & ML | Retrieval-Augmented Generation, LLM endpoints, multi-agent orchestration and intent routing, natural-language analytics, Anthropic Claude API, Model Context Protocol, Streamlit applications |
| BI & Semantic Layers | Power BI, DAX, Dremio, Denodo, SSAS Tabular and Multidimensional, SSRS, Power Apps, Power Automate |
| ERP & Enterprise Systems | SAP S/4HANA, SAP ECC FI, SAP BW, SAP AC-DOCA Universal Journal, JDE, Hyperion HFM, IBM DB2, IBM Planning Analytics, PeopleSoft, TIBCO EBX, Budget PRO |
What I bring to Databricks
For a Databricks Data Engineer role
- Production Spark and PySpark at enterprise scale, with Delta Lake design covering CDC, SCD 0/1/2, schema evolution, compaction, and Z-ORDER tuning.
- Real orchestration on the platform: Databricks Jobs (Lakeflow Jobs) coordinated with ADF and Redwood RunMyJobs, gated by Azure DevOps CI/CD with monitoring, runbooks, and hypercare.
- Databricks FinOps: cluster right-sizing and autoscaling, Delta tuning, and the Maersk managed-to-external table shift — an estimated ~$500K in annual compute savings at ADM.
- Closing the gap deliberately: Lakeflow Declarative Pipelines, Auto Loader, COPY INTO, and Asset Bundles are active exam and lab study (Databricks DE Associate → Professional in progress), tracked in this hub's Data Engineer Track.
For a Resident Solutions Architect role
- End-to-end delivery ownership: technical planning, cross-team dependencies, stakeholder alignment, and hypercare for CFO-level Finance R2R analytics.
- Architecture advisory in practice: fit-gap recommendations on batch vs. near-real-time, CDC/SCD design, and Power BI vs. Dremio/Denodo semantic-layer selection.
- Governance over sensitive data: Unity Catalog lineage plus row- and column-level controls protecting HR and finance datasets in production.
- Production GenAI on the Data Intelligence Platform: Genie, Vector Search RAG, multi-agent routing, and Databricks Apps — shipped, governed, and serving 100+ users daily.
- Mentoring and enablement: design reviews, code reviews, and PySpark standards for junior engineers.
Education
Bachelor of Technology (B.Tech), Electronics & Communication Engineering — Jawaharlal Nehru Technological University, Anantapuramu, India, 2015.