Illustration: A package with the wetransform logo on it has broken out of its cage and floats off into space suspended from a blue polygon. The cage door has been broken off and floats lower down. To the right is the text

Do You Need a Data Space?
Security, Usefulness, and Why the Biggest Risk is Not Sharing Data at All

Allgemein Interview Meinungsartikel Datenverfügbarkeit Datenmanagement Datenräume Daten Nutzwert Interoperabilität Offene Datenstandards Forschung Strategie

This summer, I was interviewed twice as a data spaces expert.

The first interview, conducted by the Ludwig-Maximilians-University Munich, focused on the role of governments when implementing sustainable digital ecosystems around data spaces and data trustees.

In this second interview, Dorothee Jahaj from the Geodateninfrastruktur Deutschland (GDI-DE) discussed ways to make sensitive data usable without compromising the interests of those who provide it with Dr. Eva Klien from the Fraunhofer IGD and myself.

“The biggest risk when sharing sensitive geodata?
That it is not shared at all.”

That somewhat provocative quote by Eva hits the sore spot directly: Focusing exclusively on the risks of re-using data will create substantial opportunity costs by preventing useful data from ever being utilised. So, how can we make sure this risk does not materialise?

This question is highly relevant in the geospatial domain. Over the past decade, Open Data initiatives such as INSPIRE have made enormous amounts of environmental and geospatial information publicly accessible. However, many of the datasets with the greatest potential to create new value are sensitive and cannot be published as Open Data.

Forest inventories, contaminated site registers, utility networks or other critical infrastructure, operational data from companies, or expensively labelled training data for AI models are just a few examples.

This is where data spaces can make a real difference, but only if we stop treating data protection as a binary choice between open and closed.

Sensitivity is Not Binary

One of the problems we repeatedly encounter is the assumption that a dataset is simply either “sensitive” or “not sensitive”. Reality is much more nuanced.

Three images side-by-side, each with a caption: "Confidentiality": Two blocks representing identities, one red and one blue. Each has a person's portrait on one visible side and bars of different length on the other. On top of each is a padlock. "Integrity": A 3x3x3 stack of blue blocks, representing a data set, with a checkmark atop each visible block. "Availability": A cloud connecting to a laptop screen, with a red downwards arrow on it

A sensible assessment should start with the potential impact of something going wrong. When evaluating protection requirements, we look at the standard dimensions:

  • Confidentiality: What happens if unauthorised parties gain access?
  • Integrity: What happens if the data is manipulated?
  • Availability: What happens if the data becomes unavailable?

In answering these questions, each concrete risk has to be identified and assessed based on evidence on the impact. Selecting effective measures that still enable the potential of the data to be leveraged is only possible on such detailed, fact-based risk assessments.

This matters because risk assessments can otherwise become dominated by diffuse concerns. Reputational damage, for example, is frequently cited as a reason to restrict access. Nobody likes to see their reputation tarnished, but this potential impact should not automatically be treated in the same way as the risks involved in exposing medical records, financial credentials, or critical infrastructure information.

There needs to be a fact-based assessment to look at plausible attack vectors, realistic consequences, and the probability that those consequences actually occur.

Our experience in the Forest Data Space illustrates this well: Forest owners usually regard forest inventory data as commercially sensitive, because they fear that greater transparency could weaken their negotiating position. Such assumptions can be tested against evidence. In countries such as Finland, comparable forest data is much more openly available than elsewhere, without obvious economic harm.

That does not mean such concerns should be dismissed out of hand, it simply means that protection requirements should be proportional to demonstrable risks.

Maximum Security, Minimal Usefulness

Sometimes, organisations that are considering becoming data providers in a data space request technical and organisational measures that make meaningful reuse of their data impossible. At that point, a data space is probably the wrong instrument for them.

A data space exists to enable data reuse. If the security requirements surrounding a dataset are so restrictive that nobody can process it in a way that creates value, adding ever more sophisticated technology on top cannot solve the fundamental issue.

We have often encountered this with anonymisation requirements. There can be an extremely narrow boundary between data that is robustly anonymised and data that has lost so much information that it is no longer useful. On top of that, even apparently anonymised geospatial information can sometimes be re-identified when combined with other datasets.

The objective should therefore not be maximum protection at any cost. It should be to maximise the intersection between protection and potential.

Access Control is Only the Beginning

Traditional information systems focus primarily on access control:

Is this person allowed to obtain this dataset?

Strong authentication, modern identity protocols, and role-based access control should be considered baseline requirements for sensitive-data environments.

However, when it comes to truly optimising the balance between protection and potential, access control alone is not sufficient. If a participant downloads a file and the data provider loses all technical control over what happens afterwards, the data space has effectively provided little more than sophisticated access management.

Data spaces introduce a second concept: usage control.

geodata, represented by a square map grid with underlying blue data layer, flows through a protected space indicated by a shield with padlock, to three entities. Industry, Research, and NGOs all receive a different sub-section of the map, which is shown beside their respective icons.

Instead of deciding only who gets the data, we define:

  • Which purpose the data may be used for
  • Which processing operation may be performed
  • Which geographical subset may be accessed
  • Whether data may be combined with other information
  • Whether data may be used for training an AI model
  • Which results may leave the controlled environment

This gives data providers considerably more confidence while simultaneously creating more possibilities for legitimate users.

The underlying principle is important:

Data sovereignty should not mean preventing data use. It should mean being able to control how data is used.

In the interview, I used the example of personal health data to illustrate this idea: Instead of handing all data from a wearable device to each application provider permanently, all of the data could remain within a trusted environment while individual analytical services receive permission to perform specific operations. This would provide the device’s owner with greater data control and security, while increasing the range of innovations they can utilise.

Bring the Algorithm to the Data

This leads to what I consider one of the most powerful architectural principles for sensitive-data ecosystems: In many cases, the raw data should not leave the data space.

Most users do not need the raw dataset, they need an answer:

  • A forest owner may need to know where storm damage is likely to have occurred.
  • An environmental authority may need an indicator derived from several sensitive sources.
  • A researcher may want to train a model, but doesn’t need to see the raw data.
  • An infrastructure operator may need an analytical result to assess cluster risks.

In all these cases, it can be much safer to bring the analytical service or AI model into the data space as a certified “smart service”, perform the processing inside this controlled environment, and return only the result.

This significantly reduces the number of situations in which raw sensitive data has to be transferred to another organisation.

“Most people do not need the data. They need the result.”

The shift from distributing datasets to providing controlled computation is one of the biggest opportunities I see for mature data spaces, and it is particularly relevant for AI.

Training and inference frequently require access to datasets that organisations would never publish openly. If models can operate close to the data while raw information remains in a trusted environment, entirely new categories of data can potentially become usable.

Data Space Infrastructure Should Become Invisible

Today, many data space demonstrations still expose the infrastructure itself to the user. A user interacts with a connector portal, initiates a contract negotiation, obtains a transfer token, and then accesses the data.

From an architectural perspective, all of this may be correct. From a user perspective, it is the wrong abstraction.

“My goal is for all this data space infrastructure to become invisible.”

A login screen set against a background of data (represented by barrel-shapes) floating in space. There's a default silhouette portrait, fields for username and password, and finally a "SIGN IN"-button.

A legitimate user should ideally authenticate once. The system should then know who they are, which credentials they hold, which policies apply, and which resources they may access.

Contract negotiation, policy evaluation, token handling, and connector communication should happen automatically behind the scenes.

For geospatial applications, this also means retaining the interfaces users already understand. Rather than forcing GIS users to learn entirely new interaction patterns, data space mechanisms should work in tandem with modern OGC APIs like OGC API Features, Maps, or Coverages.

QGIS, ArcGIS, specialised forest-management software, or analytical applications should ideally continue consuming standard geospatial services – with the data space infrastructure transparently managing trust, access and usage underneath.

This is particularly important because the real-world IT landscape is highly heterogeneous. Many public administrations still rely on GIS applications that are ten or fifteen years old and have limited support for modern identity and security protocols. Authentication and authorisation infrastructures also remain fragmented between organisations and administrative levels.

The success of data spaces will therefore depend on integrating them into the systems people actually use.

The Next Standardisation Challenge is the Data Plane

The data space community has made substantial progress on standards for discovery, identity, connectors, and automated contract negotiation.

In spite of all that, there is still a significant gap: We have spent a lot of effort defining how two parties agree that a data transaction may take place, but a far less effort has gone into standardising how the agreed policy is actually enforced during data processing and transfer.

For geospatial data, this quickly becomes domain-specific. A policy may need to express something like:

An authenticated employee of municipality X may access only those features located within the municipality's administrative boundary.

Current generic policy languages do not necessarily provide standard ways to express every such geospatial constraint. Implementations therefore start developing their own extensions, which creates a serious interoperability problem.

Trust breaks down if the same policy produces different results, depending on which provider's enforcement component happens to interpret it.

For this reason, I believe we need substantially more standardisation and certification around Policy Enforcement Points.

Different technical implementations are perfectly acceptable, but identical policies should always produce identical outcomes.

Not Every Dataset Needs a Data Space

We should resist the temptation to turn data spaces into the universal architecture for every kind of data exchange.

Data spaces are relatively sophisticated technical and organisational constructs. They require governance, trust infrastructure, participant management, agreements, connectors, policy mechanisms, and ongoing operations.

That complexity needs to be justified. Putting ordinary Open Data into a data space rarely adds much value. If a dataset can be published safely and efficiently through an open-access API, then that is usually exactly what should happen.

A data space becomes relevant where two conditions come together:

The data is sensitive and there is meaningful value in its reuse.

The additional complexity of a data space does not pay off without both sufficient value and stringent protection requirements.

Illustration of a chasm between two puzzle pieces on opposing sides. Behind the far one is a large keyhole, through which light comes in. A third puzzle piece, with a data barrel floating in space on it, hovers over the chasm and would slot in to bridge the gap perfectly.

Measure Value Through Transactions

The same principle applies economically.

“A data space only makes sense where I currently have a data gap and closing that gap creates value.”

The value created in a data ecosystem may take different forms:

  1. It improves quality
  2. It reduces costs
  3. It enables an entirely new transaction that was previously impossible

The important part is to make that value explicit. If making a dataset available enables 100 organisations to perform 1.000 valuable transactions each year, we can begin to quantify the benefit. If the resulting value is significantly higher than the cost of operating the ecosystem, there is a strong case for proceeding.

If the total value created across the sector is 24.000 € per year and the infrastructure would cost 200.000 € to set up, there probably is not.

Thinking in transactions also creates better financing models: Actors generating significant value from the ecosystem can contribute through membership fees, transaction fees, service charges, or other mechanisms. Where a substantial common-good benefit exists – as is often the case in environmental data ecosystems – the public sector should cover part of the cost as well.

On the cost side, shared infrastructure is required to make everything work. If every participant has to operate its own dedicated connector and specialised infrastructure, the economics quickly become difficult. Shared services and data trustee models can bundle hundreds or thousands of participants and dramatically reduce the cost of adding another data provider or consumer.

Communities Form Around Problems, Not Around Data Spaces

A technically excellent and economically sound data space still requires participants.

Here is a lesson we, as the data space community, have learned many times:
Most organisations do not care about data spaces. They care about solving problems.

  1. A private forest owner may care about deciding how to adapt a forest to climate change.
  2. A forest-management software company may care about giving its customers access to additional data sources.
  3. An authority may care about making a funding programme more effective.

In the Forest Data Space, for example, services can provide forest owners with information about disturbances detected from frequently updated data. A message saying “Something may have happened in your forest, take a look!” has an immediately understandable benefit.

The underlying data space architecture enables the service, but it is not the reason anybody uses it. This distinction is perhaps one of the most important lessons for the next generation of data ecosystems:

Lead with the problem and the value-added service. Let the data space disappear into the infrastructure underneath.

From Protecting Data to Enabling Trusted Use

For a long time, the discussion around sensitive data has been dominated by a defensive question:

How do we prevent anything from going wrong?

That question remains important. However, to be competitive, to enable innovation, and to scale our digital economy, we need to ask the following question with the same urgency:

What value are we losing because useful data cannot currently be accessed?

Mature data spaces can help us move beyond the false choice between publishing everything and sharing nothing.

They can combine proportionate protection, trusted identities, standardised policies, usage control, data-proximate processing, and familiar APIs into an infrastructure that makes previously inaccessible data usable.

Achieving that requires a change in mindset: The goal of assessing risks in this context should not be maximum restriction, it should be maximum useful and trustworthy reuse within an acceptable level of risk.

If we get that balance right, data spaces can become a critical layer of infrastructure that can turn data that is currently locked away into a resource for better decisions, new services, more efficient processes, and trustworthy AI.


Should you wish to learn more about whether data spaces are the right solution for you, we'd be happy to assist! Reach out here.