Skip to content
Ian Cunningham monogramIan CunninghamData & AI consultant

Blog

Start with Decisions, Not Tables: Planning a Data Warehouse

A practical guide to defining the decisions, business processes, ownership, scope, and delivery approach that should shape a data warehouse before technology choices begin.

Start with Decisions, Not Tables: Planning a Data Warehouse
KT

Article summary

Key Takeaways

  1. A warehouse plan begins with decisions and processes

    Source systems and tables matter, but they don't define the analytical outcome or its business value.

  2. Business requirements must meet data reality

    Grain, history, quality, ownership, freshness, and operational constraints determine whether a requested answer is supportable.

  3. Inmon and Kimball provide useful emphases, not a compulsory binary choice

    One highlights enterprise integration; the other shows how coordinated dimensional increments can deliver around business processes.

  4. Incremental delivery still needs an integration strategy

    Without conformed definitions, ownership, and architectural guardrails, a sequence of fast releases can create another fragmented estate.

  5. Success includes trust and operability

    Loading data is only one part of delivering an analytical service people can understand, validate, secure, and support.

A warehouse project can appear well organised before anyone has agreed what it needs to achieve.

There may be a platform decision, a source-system inventory, a delivery partner, and a roadmap full of ingestion work. The project can tell you where the data lives and how it might be moved, but still struggle with more basic questions:

  • Which decisions should become easier or more reliable?
  • Which measures need to agree across departments?
  • What history needs to be retained?
  • Who can approve a business definition?
  • Which business process should be delivered first?
  • Who will operate the resulting analytical service?

Those aren’t questions to postpone until reporting begins. They should shape the scope, data model, controls, delivery sequence, and acceptance criteria from the start.

A Source List Isn’t a Warehouse Plan

A source inventory is useful. It identifies systems, owners, interfaces, formats, volumes, and extraction constraints that the team will eventually need to address, but it doesn’t establish the analytical purpose.

A list containing a CRM system, finance application, ecommerce platform, support tool, and several spreadsheets can’t answer questions such as:

  • Which business process are we trying to understand?
  • Which decisions or actions should the analysis support?
  • What would each row of analytical data represent?
  • How current does the information need to be?
  • Which corrections should change previously reported results?
  • Who owns the definitions and quality thresholds?

Start with Decisions and Business Processes

Suppose an organisation wants better visibility of sales and fulfilment.

That sounds reasonable, but it leaves several interpretations open. A more useful planning conversation connects the desired outcome to a business process and the decisions made within it.

Sales and fulfilment provide one concrete example below. The same planning sequence could begin with reducing claim-settlement delays, understanding loan arrears, improving patient flow, or resolving support cases more consistently.

  1. 1
    Goal

    Improve visibility of sales and fulfilment.

  2. 2
    Decision

    Identify where operational leaders should intervene when orders are delayed.

  3. 3
    Business process

    Follow an order from placement through allocation, dispatch, and delivery.

  4. 4
    Potential measures

    Orders by stage, elapsed time, late-order rate, and cancellation rate.

  5. 5
    Descriptive context

    Customer, product, location, channel, fulfilment route, and date.

  6. 6
    Historical need

    Retain the statuses and responsibilities required to explain past performance accurately.

The subject can change without changing the reasoning: move from a broad goal to a decision, identify the process and available evidence, then establish measures, context, history, ownership, and an achievable delivery boundary.

The measures and descriptive contexts are provisional at this point. Source investigation may show that a requested event isn’t captured, a timestamp has a different meaning, or historical statuses have already been overwritten.

Good requirements work moves between business need and data reality. It doesn’t assume that every valuable question can be answered from the data currently available.

What the Warehouse Requirements Need to Cover

Functional requirements alone don’t describe a dependable analytical service. A project also needs expectations for quality, history, security, recovery, support, and change.

Area Questions to establish
Decisions and outcomes Who will act differently, and what evidence will they need?
Business meaning Which processes, events, entities, measures, and definitions are involved?
Data reality Which sources capture them, at what grain, with what quality and history?
Service expectations How fresh, available, responsive, and recoverable must the data be?
Control Who may access it, who owns definitions, and what must be auditable?
Delivery What is the smallest coherent release, and which dependencies could prevent it?
Operation Who monitors, supports, changes, and pays for the service after launch?

These areas affect one another. A requirement for historically accurate customer segmentation affects the model, source assessment, transformation logic, storage, testing, and possibly the business process used to approve customer changes.

Inmon and Kimball as Strategic Lenses

Discussions about warehouse planning often introduce Inmon and Kimball as opposing methodologies. The familiar shorthand is Inmon as top-down and Kimball as bottom-up.

That comparison is useful, but only if we avoid turning it into a slogan.

Inmon and the enterprise-first perspective

Bill Inmon’s work begins with a broadly integrated enterprise foundation. Operational data is organised around major subjects such as customer, product, and finance, then used to supply departmental data marts and other analytical services. The wider collection of warehouse, marts, operational data stores, metadata, and supporting components is known as the Corporate Information Factory.

The enterprise-first emphasis encourages teams to ask:

  • Which subjects and definitions need to be consistent across the organisation?
  • Where do business definitions cross departmental or operational boundaries?
  • What should form the durable integrated foundation?
  • How will departmental marts depend on the central warehouse?
  • How will metadata, history, lineage, and ownership be governed?

This can demand substantial coordination. Departments may need to agree definitions before their immediate reporting requirement is delivered, and the programme needs enough organisational participation to sustain the enterprise scope.

It would be misleading to say that every Inmon-influenced programme must deliver one enormous warehouse before anybody receives value. The more useful point is that enterprise integration shapes the foundation rather than emerging only from a collection of local analytical models.

Bill Inmon’s Data Warehousing 2.0 paper describes how ETL, departmental marts, operational data stores, metadata, and the Corporate Information Factory grew around the original warehouse concept.

Kimball and the business-process perspective

Ralph Kimball’s approach starts with a prioritised business process, such as orders, deliveries, or claims, and delivers an analytical model around it. Later models remain compatible by reusing agreed descriptions of shared concepts such as date, customer, and product. Dimensional modelling calls these reusable structures conformed dimensions.

Kimball’s four-step dimensional design process asks teams to:

  1. 1Select the business process
  2. 2Declare the grain
  3. 3Identify the dimensions
  4. 4Identify the facts

Article 4 examines those decisions in detail. At the planning level, the approach encourages a useful analytical increment that can later connect coherently with other processes.

The word bottom-up can create the wrong impression if it suggests that departments are free to build unrelated star schemas. Kimball integration depends on coordination. Shared dimensions such as date, customer, product, or organisation need compatible meanings if users are expected to analyse several business processes together.

Kimball calls this coordinated process-by-process approach the enterprise data warehouse bus architecture. A planning tool called the bus matrix places business processes in rows and shared dimensions in columns, making it easier to see where definitions must be reused.

The Kimball Group’s dimensional modelling guidance explicitly includes business requirements, data realities, collaborative modelling, conformed dimensions, and the enterprise bus architecture. Fast departmental isolation isn’t the goal.

What the Approaches Emphasise

The comparison is better understood as different sequencing and integration emphases.

Enterprise-first emphasis means establishing a broadly integrated organisational foundation before supplying more focused analytical structures. In this article, business-process-first emphasis means beginning with a prioritised process that can produce a useful analytical model, then expanding through coordinated models and shared definitions.

Planning concern Enterprise-first emphasis Business-process-first emphasis
Initial scope Broad integration model and shared foundation Prioritised analytical process or subject area
Early delivery More dependency on cross-domain agreement Earlier value when the first process is genuinely useful
Integration mechanism Enterprise model and governed central integration Shared dimensions and coordinated process models
Organisational demand Sustained enterprise participation and architecture capacity Strong prioritisation and discipline across increments
Main failure risk Long planning effort without visible use Fast delivery that fragments when conformance is neglected

These are tendencies rather than laws. Real implementations vary, and either approach can be weakened by unclear ownership, poor source understanding, or a delivery plan that the organisation can’t sustain.

A Pragmatic Iterative Position

Many programmes combine the underlying ideas:

  • Establish enterprise principles for ownership, naming, security, metadata, and integration.
  • Select a valuable business process for the first coherent release.
  • Design its analytical model collaboratively with business and technical participants.
  • Identify the definitions and dimensions that future processes will need to share.
  • Retain room for other data products and analytical workloads.
  • Learn from delivered use before expanding the scope.

This isn’t a new methodology with a universally correct recipe. It’s a way of making the integration and delivery decisions explicit.

Operationally, the three planning emphases can be described as follows:

  • Enterprise-first path: Establishes a broad integrated foundation before supplying dependent analytical structures.
  • Business-process-first path: Delivers coordinated dimensional increments connected through shared definitions.
  • Pragmatic iterative path: Establishes enterprise guardrails while delivering one prioritised process, then expands using evidence from the working service.
Three warehouse delivery paths compare an enterprise-wide foundation, coordinated business-process increments, and an iterative approach that combines an initial release with shared architectural guardrails

The paths describe planning emphases rather than compulsory or mutually exclusive methods.

Microsoft’s current dimensional modelling guidance for Fabric Warehouse makes a similar practical recommendation: treat the enterprise warehouse as an important foundation, but build it iteratively by starting with the most important subject areas.

Select the First Release Carefully

The first release should be small enough to deliver and substantial enough to be useful. A convenient dataset isn’t necessarily a coherent analytical outcome.

  • It supports an identifiable decision and named users.
  • The required events and descriptive data are available at a usable grain.
  • Important definitions can be agreed within the release boundary.
  • Cross-domain dependencies are understood rather than hidden.
  • Some dimensions or measures are likely to be reusable later.
  • The result can be reconciled and validated independently.
  • The operating team can monitor and support it after launch.

I wouldn’t turn those considerations into a simplistic score where every factor has the same weight. Some are gates. A commercially valuable use case may still be unsuitable if the event it depends on isn’t captured reliably.

Define Success Beyond the Data Load

A technically successful pipeline can load every morning while the warehouse remains difficult to trust or use.

More useful acceptance criteria might include:

  • Agreed measures reconcile to authoritative sources within defined tolerances.
  • Historical reports remain stable under the agreed correction policy.
  • Named users can answer priority questions without rebuilding core logic.
  • Refresh and recovery meet the agreed service expectations.
  • Access rules have been tested using representative roles.
  • Definitions, lineage, and known limitations are discoverable.
  • The owning team can monitor, support, and change the service.

Dashboard usage on its own doesn’t prove that decisions have improved. Equally, a successful load doesn’t prove that the data is understood, reconciled, or useful.

Produce Decisions and Artifacts, Not Only a Requirements Document

A requirements document may be part of the output, but planning should leave behind evidence that can guide design, delivery, testing, and operation.

  • Problem and outcome statement
  • Stakeholder and decision map
  • Business process definition
  • Metric and terminology register
  • Source and data-quality assessment
  • History and correction requirements
  • Security and governance responsibilities
  • Service expectations
  • Initial release and dependency map
  • Architecture decision records
  • Validation and acceptance approach
  • Delivery roadmap and operating ownership

These artifacts make planning decisions easier to find and apply. If important decisions exist only in meeting notes or individual memory, later technical work will fill the gaps with assumptions.

Questions to Take into Architecture Design

The plan now needs to become a source-to-consumption design.

  • What must be captured from each source, and can extraction be repeated safely?
  • Where will validation and standardisation occur?
  • Which data should remain close to its source representation?
  • Which structures will serve reporting and analysis?
  • How will historical change and corrections flow through the layers?
  • Where do security, lineage, orchestration, and monitoring apply?

Those questions lead into the next article, which examines the layers between operational sources and usable analytical information, including ingestion, staging, transformation, storage, presentation, governance, and observability.

Expanded diagram

Work with Ian

Need help turning a complex data or technology requirement into something workable?

If this post connects with a problem you are facing, I can help clarify the requirement, shape the approach, and move it toward a practical solution.