Skip to content
sabesp-transforms-operations-with-lakehouse

Sabesp Transforms Operations with Lakehouse

Sabesp (Companhia de Saneamento Básico do Estado de São Paulo) supplies drinking water and provides wastewater collection and treatment across 376 municipalities in the state of São Paulo serving around 30 million people. It is one of the world’s largest water and sanitation utilities, and the largest in Brazil. 

Sabesp aims to leap five decades in five years, expanding access to safe water and sanitation for millions. The company is committed to bringing forward by four years the targets set under Brazil’s Sanitation Legal Framework, with the goal of delivering dignity, public health, and sustainable development while safeguarding natural resources for future generations. 

As regulatory requirements evolved and customer expectations increased, the company embarked on a broad digital transformation to improve operational efficiency, strengthen customer experience, and enable faster, more informed decision-making across the business. To support this transformation, Sabesp modernized its legacy Oracle Exadata-based data warehouse and adopted Databricks as the foundation for its data and AI strategy. Today, Databricks powers a Data Mesh architecture spanning more than 23 business domains, supports large-scale SAP data ingestion, and enables self-service analytics and AI-driven applications across the organization.

Building a Modern Data Foundation for a New Era of Water Utilities

The water and sanitation industry in Brazil is undergoing significant change. New legal and regulatory requirements and increasing operational complexity have created pressure to improve efficiency while continuing to deliver reliable service to millions of customers.

For Sabesp, data became a strategic asset in navigating this transformation. The organization needed a platform capable of supporting operational analytics, customer service modernization, fraud detection, and AI-driven innovation across a diverse set of business functions.

Prior to Databricks, however, the company’s data architecture was built around a traditional Oracle Exadata data warehouse. While the platform had served the organization for years, it became increasingly difficult to scale as data volumes and analytical demands grew.

The legacy environment created several challenges. Infrastructure costs continued to rise, data remained difficult to access across business functions, and complex ETL processes slowed the delivery of new use cases. At the same time, the architecture was not designed to support modern AI, machine learning, or real-time analytics workloads.

Enabling Data Mesh Across 23+ Business Domains

As Sabesp evaluated modernization options, the team sought more than a replacement for its data warehouse. The goal was to create an operating model that would allow business domains to take ownership of their own data while maintaining governance and consistency across the organization.

Databricks stood out because it unified data engineering, analytics, governance, and AI on a single platform. The lakehouse architecture eliminated the separation between data lakes and warehouses, while Unity Catalog provided the governance foundation required to support decentralized ownership.

The migration from Oracle Exadata became the catalyst for implementing a Data Mesh architecture across the company. Rather than relying on a centralized team to manage every analytical request, business domains could begin developing and managing their own data products.

“The migration from Exadata to Databricks enabled Sabesp to implement a Data Mesh architecture — something that was not feasible on the legacy platform,” said Eric Leite, Chief Data Officer. “This decentralized approach empowers each of the 23+ business domains to own and deliver their data products independently, dramatically accelerating time-to-value and reducing dependencies on a central data team.”

To support this model at scale, Sabesp standardized development and deployment processes using Databricks Asset Bundles (DABs), creating consistent CI/CD workflows across all participating domains. The result is a platform that combines local ownership with centralized governance, allowing teams to move faster without sacrificing control.

Accelerating SAP Analytics and AI-Powered Data Access

Ensuring accessible data is a fundamental requirement for an organization of Sabesp’s scale and complexity. The company manages information spanning customer billing, consumption, geography, field operations, and infrastructure systems, alongside significant volumes of SAP data. Databricks’ BDC Connector became a key component of the modernization effort, enabling large-scale ingestion of SAP data into the lakehouse and creating a more complete foundation for analytics. 

At the same time, the organization has focused on making data more accessible to business users. Sabesp uses Databricks Genie and AI/BI Dashboards to enable self-service analytics across the enterprise.

Genie has been particularly impactful because it allows users to interact with data using natural language rather than SQL. Business analysts and operational managers can ask questions in Portuguese and receive insights directly, removing technical barriers that previously limited access to analytics.

“Sabesp is leveraging Databricks AI/BI dashboards and Genie for natural language data exploration. Genie has been particularly transformative — enabling business analysts and operational managers to query data in Portuguese using natural language, without needing SQL expertise. This has democratized data access across the organization, especially for non-technical business domain owners,” said Leite.

The company is also using Databricks Assistant to accelerate development, debugging, and SQL optimization, helping teams onboard to the new platform more quickly as the migration continues.

Creating a Foundation for AI, Scale, and Future Growth

Today, Databricks serves as the foundation for Sabesp’s long-term data and AI strategy. The platform provides the scalability required to support the needs of approximately 30 million customers while enabling new capabilities that were difficult or impossible to implement in the legacy environment.

Serverless SQL allows the organization to scale compute resources cost-efficiently, on demand, across multiple business domains, without managing infrastructure. Unity Catalog provides centralized governance and fine-grained access controls across the Data Mesh. Foundation Model APIs give teams direct access to leading large language models for AI applications without requiring additional infrastructure investments.

These capabilities are already supporting initiatives such as AI-powered customer service experiences, fraud detection solutions, and natural-language analytics. More importantly, they provide a framework for continued innovation as business needs evolve.

For Sabesp, the migration was more than a technology modernization project. It established a scalable foundation that enables data ownership, self-service analytics, and AI innovation across one of the largest utility organizations in Latin America.
 

colind88

Back To Top