Explainer
    Concepts
    Concepts

    Part of The company brain: where your organization's knowledge lives

    What is data architecture?

    Data architecture is the layout that decides where data arrives, where meaning is recorded and where answers come from. Four layers, and the choices that recur.

    DatahubDatahub editorial team8 min read
    Share on
    What is data architecture?

    Data architecture is the layout of your data landscape: where data arrives, where it is stored, where meaning is recorded, and where answers come from. It is the floor plan every later choice lands on.

    You do not recognise good architecture by the number of components, but by being able to point at where every number came from.

    The four layers

    05An answer the meeting trustsA report, a process or an agent04MeaningDefinitions, owners, agreements03ModellingFrom raw source to usable shape02Storage and computeLakehouse, storage, processing01SourcesThe systems where the work happensFrom raw source to answer. Each layer does one thing.

    Source and ingestion. Systems deliver: ERP, WMS, TMS, sensors, spreadsheets. Nothing is changed here, only recorded as it stood and when. That timestamp matters more than it seems: without it you cannot later reconstruct why a number was different last week.

    Storage. One place where raw and processed data sit side by side, usually a lakehouse. The point is not the technology but that there is one place instead of four.

    Meaning. Definitions, lineage, quality rules, access. This is the layer most organisations skip and the one that decides whether the rest can be trusted.

    Use. Reporting, planning, models, agents. What happens here is only as good as the layer underneath.

    The choices that keep coming back

    A DRAWING ON A SLIDEARCHITECTURE THAT HOLDSStarts with toolingStarts with the questionMeaning comes laterMeaning is a layerNo one owns a termEvery term has an ownerEvery new source is a projectEvery new source plugs inThere is no universally right answer, but there is an explicit one.

    Batch or streaming. Central or per domain. Copy or query at the source. One catalogue or one per department. These choices are not right or wrong, but they must be deliberate and written down. Architecture that emerges per project is not architecture, it is sediment.

    Four questions sharpen every choice:

    • Who asks the question, and how often? A number that lands in a monthly meeting asks for something different than a signal a planner needs every fifteen minutes.
    • How fresh must the answer be? Freshness is the most expensive requirement in any architecture. Ask for it only where it genuinely changes the decision.
    • Who is allowed to see it? Permissions arranged only in the reporting layer are a leak. They belong in the meaning layer.
    • What must you be able to explain later? When an auditor or a customer asks how a number came about, traceability is not an extra: it is the point.

    Four patterns you actually meet

    Almost every architecture is a variant of four shapes.

    Central data warehouse. Everything is modelled into one schema. Predictable and easy to control, but every new question costs modelling work, and semi-structured sources fit badly.

    Data lake. Everything lands as it is. Cheap and flexible, but without agreements it turns into an archive nobody trusts.

    Lakehouse. Open storage with transactions, schema and a catalogue on top. In practice the current starting point for most organisations: one place for analytics, engineering and AI. What that means concretely is covered in what is Databricks.

    Federated or mesh. Domains deliver themselves, on one shared platform. This is an organisational choice on top of a lakehouse, not instead of one. See what is a data mesh.

    Most organisations sit on a hybrid, and that is fine. It only becomes a problem when nobody can write down which shape applies where.

    Meaning is not a by-product

    The most expensive mistake is assuming meaning appears once the technology stands. It does not. Without recorded definitions you get a technically immaculate platform in which two departments still report different revenue. That is why the meaning layer belongs in the design, not in the aftercare. How to record it per dataset is covered in what is a data contract.

    Practically, that layer consists of four things you can point at: a list of terms and their definitions, an owner per term, lineage per field back to the source system, and role-based permissions enforced in one place.

    The requirements that are not a layer

    Besides layers, an architecture has requirements that cut across all of them, and you should fix them in numbers up front:

    • Freshness. Per dataset: how old the answer may be.
    • Recovery. How long a delivery may stall, and how far back you can roll.
    • Cost. What a question may cost, and who sees that bill.
    • Access. Which role sees which row, and where that is checked.
    • Retention. What stays, for how long, and how deletion becomes provable. See effective data lifecycle management.

    Without numbers these are opinions, and opinions cannot be tested at delivery.

    Where to start

    01The questionWhich decision depends on it02The answerWhich number, in which unit03The definitionWhat counts, what does not04The sourceWhich system, which field, how fresh05The gapWhat is missing here is your first decisionDrawing one question all the way back beats a picture of the end state.

    Not with a tool and not with a drawing of the end state. Start with one question the business asks every month and draw the route back: which answer, which definition, which source, how fresh. The gaps you find there are your first architectural decision.

    Do that for three questions from different corners of the organisation and you have no diagram, but a list of decisions with reasons attached. That is the only form of architecture still standing a year later.

    How to write it down

    Keep the document small enough to maintain. What works in practice: one page per layer with the chosen shape and why, one table with the requirements in numbers, and one register of terms, owners and sources. Anything larger than that does not get maintained and is fiction within two quarters.

    If you want to see how we build and run those layers, look at the engine. If you want to assign ownership per domain, read what is a data mesh.

    About the publisher

    Datahub

    Datahub editorial team

    Pieces without a personal byline are written and reviewed by the Datahub team. We build governed data foundations for logistics, retail and manufacturing, and only publish figures we measured ourselves or read in a primary source.

    Why this source

    • Every publication is reviewed before it goes live
    • Figures follow the methodology at /research/methodology

    Writes about: Data foundations · AI readiness

    More about the team

    Next step

    Want to see what's already inside your organization?

    Leave your details. We'll reach out and plan a scan. Within thirty days you'll see one concrete result.

    No newsletter, no reselling. Just this conversation.

    Comments

    Comments are reviewed by the editors before they appear.

    Use your Google or Apple account, or your business email address.

    Sign in to comment