Data warehouse
A data warehouse (DWH) is a central database system optimised for analysis that brings data from different sources together, unifies it and stores it long term. Unlike operational databases, designed for fast individual transactions, a data warehouse is designed to analyse large amounts of historical data efficiently. It thus forms the basis for business intelligence, marketing dashboards and data-driven decisions. Well-known cloud solutions are Google BigQuery, Snowflake, Amazon Redshift and Microsoft Fabric.
Also known as: data warehouse, DWH
What is a data warehouse?
In most companies, data is scattered: in the CRM, in the shop system, in accounting, in Google Analytics 4 and in various marketing tools. Each of these systems has its own view of the world, its own labels and its own data formats. As long as this data stays separate, overarching questions such as a customer’s true value across all channels can hardly be answered.
A data warehouse solves that problem by serving as a central collection point. Data from every relevant source is imported regularly, converted into one uniform model and stored permanently. On that consolidated basis, analysis becomes possible that individual tools could never deliver.
The decisive difference from an ordinary database lies in the purpose. An operational database serves day-to-day running and has to process many small writes quickly. A data warehouse is optimised for the opposite: reading and aggregating vast amounts of data over long periods, as reporting and analysis typically require.
How data gets into the warehouse
The data's route into the warehouse follows a process traditionally called ETL: extract, transform, load. Data is first extracted from the source systems, then transformed, that is, cleaned and unified, and finally loaded into the warehouse. In modern cloud architectures the ELT variant has often prevailed, in which the raw data is loaded first and transformed only in the warehouse.
Specialised tools are often used to connect the sources. Services such as Fivetran or Airbyte handle automatic transfer from hundreds of systems, while tools such as dbt organise the transformation of the data within the warehouse. That creates a repeatable, documented data pipeline.
Clean modelling matters. Data from different sources has to be structured so metrics are unambiguously defined and the same terms mean the same thing everywhere. Without that shared data model, contradictory figures arise and undermine trust in the whole reporting.
Cloud data warehouses compared
The market is shaped by a few large cloud platforms differing in architecture, pricing model and ecosystem. Google BigQuery meshes closely with the Google marketing stack and is particularly popular for analysing raw Google Analytics 4 data. Snowflake counts as platform-neutral and flexibly scalable. Amazon Redshift is deeply anchored in the AWS ecosystem, and Microsoft Fabric bundles several data services into one platform.
The overview below classifies the four solutions by the most important decision criteria. For GDPR compliance, choosing an EU region matters in particular.
| Platform | Strength | scaling | EU region | Pricing model |
|---|---|---|---|---|
| Google BigQuery | GA4 raw data, close to marketing | Serverless, automatic | Yes (e.g. europe-west) | Per query / per storage |
| Snowflake | Platform-neutral, flexible | Compute scales separately | Yes (EU regions) | Per compute time |
| Amazon Redshift | AWS integration | Cluster and serverless mode | Yes (EU regions) | Per cluster / per use |
| Microsoft Fabric | All-in-one, close to Power BI | Capacity-based | Yes (EU regions) | Capacity plan |
Data warehouse and data protection
As soon as personal data flows together in a data warehouse, the GDPR's strict requirements apply. The choice of storage location is decisive: if an EU region is used, the data stays within the European legal area, which avoids transfers to third countries. All the platforms named offer EU regions, but the configuration has to be set deliberately.
Data minimisation also plays a central role. Not all raw data has to be stored unchanged. Pseudonymisation, aggregation and defined deletion periods reduce the risk and meet the principle of storage limitation. A well-considered permissions concept also ensures only authorised people access sensitive data.
At Elisabit we design data warehouse setups to be technically solid and privacy-compliant at the same time. Data from first-party sources can thus be brought together sensibly without first-party data becoming a legal risk.
- 1Identify and prioritise the relevant data sources
- 2Set an EU region for the warehouse and review data transfers
- 3Build and automate the data pipeline via ETL or ELT
- 4Define a unified data model with clear metrics
- 5Implement permissions, pseudonymisation and deletion periods
- 6Data in BI tools and Marketing dashboards make available
When a data warehouse pays off
A data warehouse pays off as soon as you regularly have to merge data from several systems to answer questions no single tool covers. Typical triggers are cross-channel Attribution, a unified view of the customer, or the need for reliable, automated reporting.
For smaller companies with manageable data volumes, a full data warehouse can at first be oversized. Thanks to serverless cloud models and usage-based billing, though, the barriers to entry have fallen: you largely pay only for the data and queries actually processed, which makes starting out economical for mid-sized companies too.
The real value only emerges through the connection to analysis and visualisation. A data warehouse is not an end in itself but the basis on which BI tools and Looker Studio build reliable reports. Only that combination turns collected data into a real basis for decisions.
Frequently asked questions
What is a data warehouse?
A data warehouse is a central database system optimised for analysis that brings data from many sources together, unifies it and stores it long term. Unlike operational databases it is designed to analyse large amounts of historical data efficiently. It forms the basis for business intelligence, reporting and data-driven decisions.
How does a data warehouse differ from an ordinary database?
An operational database serves day-to-day running and has to process many small writes quickly. A data warehouse is optimised for the opposite: reading and aggregating vast amounts of data over long periods. It also stores historical data permanently, while operational systems mostly hold only the current state.
Which data warehouse is the right one?
The choice depends on your ecosystem. Google BigQuery is ideal if you want to analyse Google Analytics 4 raw data. Snowflake convinces through platform neutrality, Amazon Redshift through AWS integration and Microsoft Fabric through its closeness to Power BI. For the GDPR, choosing an EU region is decisive with all of them.
Is a data warehouse GDPR compliant?
A data warehouse can be run in a GDPR-compliant way if personal data is stored in an EU region, transfers to third countries are avoided and the principles of data minimisation are observed. Pseudonymisation, defined deletion periods and a clean permissions concept are central building blocks here. The configuration has to be set deliberately to match.
What do ETL and ELT mean?
ETL stands for extract, transform, load: data is extracted, cleaned and then loaded into the warehouse. ELT reverses the order and loads the raw data first, transforming it within the warehouse. ELT is widespread in modern cloud architectures, because the warehouse's compute handles the transformation efficiently.
Is a data warehouse worth it for small companies?
Thanks to serverless cloud models and usage-based billing, starting out is economical for mid-sized companies too, since you largely pay only for data actually processed. A data warehouse pays off as soon as you regularly have to bring data from several systems together. With very small amounts of data, though, a full warehouse can at first be oversized.
Related terms
First-party data is data a company collects directly from its own users, with their consent.
A customer data platform (CDP) brings customer data from every source together into one central, unified profile and makes it available to other systems.
Software that brings data from different sources together, evaluates it and visualises it in interactive dashboards.
A free Google tool for data visualisation that brings GA4, Google Ads, Sheets and more together in dashboards.
Collecting, measuring and analysing website data in order to improve content and conversions in a targeted way.
Tools that manage tracking and marketing tags centrally, without having to change the website code for every adjustment.
Put AI to work for your business?
We help you integrate artificial intelligence into your processes, your marketing and your website — strategically and securely.

Your contact
Stefan
I look forward to hearing about your project and finding the best solution together.