Friday, July 10, 2026

Dive into J-Spaces: Anthropic’s New Research on AI’s Inner Workspace

Anthropic published this week an interesting research paper about LLMs. They may have found something in large language models that looks a lot like a working space for thought. Not that we say "AI is consious" or it has a mind, but rather something more subtle.

Global Workspace LLM
source: https://transformer-circuits.pub/2026/workspace/index.html

Inside a model, there seems to be a small privileged space where certain concepts become available for the model to work with. Anthropic calls this J-space (a model representational space). You can think of it as something like a temporary workbench: a place where intermediate ideas, judgments, and partial conclusions show up before the model gives its final answer.

That matters because we usually judge AI from the outside. Normally we are looking for the answers like: is this correct, safe etc.

But the harder question is what happened before the answer appeared. Was there any reasoning? Something suspicious happened? Not disclosed mistake? Maybe even right answer but for the wrong internal reason? Until now, LLM interaction was some kind of a blackbox that has an input and produces a text outside. This research gives us a way to start looking at some of that hidden process.

One example I like is arithmetic. A model might only output the final answer, but internally you can sometimes see intermediate steps appear before the final response. In other cases, the model may recognize that a piece of code contains a bug, or that a prompt looks suspicious, before it says anything about it.

source: https://transformer-circuits.pub/2026/workspace/index.html

 That has big implications for AI safety. The future of trustworthy AI will not just be about making models more polite or better at refusing bad requests. It will be about understanding what is happening inside them. Because an answer can look fine on the surface while the internal reasoning behind it is not fine at all.

That is the part I find most important here. We are moving from “Does the model say the right thing?” toward “Can we inspect how the model got there?” There is still a lot to be careful about. This is not proof that AI thinks like humans. But it does suggest that these models may have internal structures that behave like a working memory for concepts.

And if we can observe that space, even imperfectly, we may get better at detecting hidden reasoning, strategic behavior, mistakes, and mismatches between what a model says and what it is internally representing.

To me, that feels like a meaningful step. Now we can make AI systems not only more capable, but more inspectable.


Link to full article: https://transformer-circuits.pub/2026/workspace/index.html

External commentary: https://www-cdn.anthropic.com/files/4zrzovbb/website/cc4be2488d65e54a6ed06492f8968398ddc18ebe.pdf

 

 

Sunday, May 24, 2026

From Farms to Blockchains: Understanding Ledgers, Daml and Canton

I often see terms like blockchain, ledger, Daml, smart contracts, and imagine something highly abstract and technical. Now I want to explain using my favorite analogy - farming.

Imagine several farms:

  • Your farm
  • John's farm
  • Peter's farm
  • A dairy cooperative
  • A bank

They all do business with each other: selling cows, leasing land, financing equipment, making agreements.

Everyone needs to keep track of who owns what.
A traditional database is like a barn full of paper records.

Inside you can write things like:

  • John owns 5 cows
  • Peter sold 2 pigs
  • Mark leased a field
The barn stores information well. But the barn itself doesn't enforce rules. It doesn't know:
  • Can John sell the same cow twice?
  • Can Peter see Mark's private agreement?
  • What happens if two people try to sell the same cow at the same time?
  • Who is allowed to change records?
  • What happened first?
A database stores information, but business rules usually need to be enforced separately by applications and processes.

Daml 

This is where Daml enters.

Daml lets you describe the rules of the business:

To sell a cow:
  • the seller must own it
  • both parties must approve
  • after the sale, the previous owner no longer owns it
Daml defines the business rules and workflows that participants must follow.

Ledger

Think of the ledger as a farm manager. A manager who not only checks the rules but also maintains the official state of who owns what.

The manager checks:

  • Are the rules followed?
  • Is the cow really owned by the seller?
  • Has it already been sold?
  • Who can see this transaction?

So the system becomes:

Farmer - submits agreement
Daml - defines rules
Ledger - enforces rules, manager
Database - stores records

Blockchain

Now imagine there is no single manager. Instead, every farm keeps its own notebook.

When a cow is sold:

  1. You write: "I sold a cow to John."
  2. Other farms in the network verify:
    • Was the cow yours?
    • Does it exist?
    • Was it already sold?
  3. If everyone agrees, everyone updates their notebook.

That is essentially a blockchain.

Example: public blockchain (Bitcoin, Ethereum)

Anyone can join, think: "Any farmer in the world can keep a copy of the notebook." Thousands of participants maintain and validate the same system. Everyone sees everything.

Enterprise ledger (Daml + Canton)

Think of a private agricultural cooperative, only trusted members participate:

  • Bank A
  • Bank B
  • Exchange
  • Insurance company

Each participant can run their own machine, but random people on the internet cannot simply join.

And unlike public blockchains, not everyone sees everything. John only sees agreements involving John. Banks only see transactions they are allowed to see. Privacy is built into the model.

So far we have:

  • Blockchain - thousands of farmers sharing copies of the same notebook
  • Ledger - a system enforcing rules and maintaining records
  • Daml - the language defining the rules
  • Database - where the papers are stored

Every blockchain is a ledger.
Not every ledger is a blockchain.

Canton 

Canton Network is a distributed ledger network that shares some characteristics with blockchains, but works differently from Bitcoin or Ethereum. Instead of showing every transaction to everyone, it is built for privacy. Think of it as a "network of networks."

Only the people involved in a transaction can see its details. Other participants cannot. This is called a "need-to-know" model — you only see the information you actually need, not everything happening on the network.

Smart Contracts 

In Daml, business rules are implemented as smart contracts — digital agreements with built-in logic that automatically execute when conditions are met.

Think of it like farming:

Imagine a farmer, a grain buyer, and a transport company making a deal.

The agreement says:

  • the farmer delivers 100 tons of wheat,
  • the transport company confirms delivery,
  • the buyer sends payment.

With paper contracts, people have to check everything manually. With a Daml smart contract, the rules are already written into the system. Once delivery is confirmed, the contract can automatically move to the next permitted step, such as triggering the payment process.

The important difference is that not everyone sees the contract. Only the parties involved can access it.

Putting everything together:

  • Database stores records
  • Daml defines business rules
  • Smart contracts apply these rules automatically
  • Ledger validates and maintains the shared state
  • Blockchain is one way of running a ledger across many independent participants
  • Canton focuses on shared records with privacy and selective visibility between participants.

 

 

 

Wednesday, May 20, 2026

Qdrant is staying in my setup for good.

Recently, I moved from application-level cosine similarity comparisons to a dedicated vector search engine running at the storage layer. Qdrant significantly improved the performance of semantic search in my pipeline, reducing similarity search time from around 40 seconds to roughly 1 second for the same workload.

Qdrant logo
What surprised me most was not only the raw speedup, but how much architectural complexity disappeared after the migration.

The previous approach was simple: embeddings were stored as regular data and compared inside the application layer. That was a reasonable starting point.

But I was indexing Flink, which is a big Repo, the architecture crossed a clear boundary.
 
At that point, the application was doing work that belonged to a specialized vector search engine:
- loading large embedding sets into memory
- calculating similarity scores one by one
- sorting candidates in the application process

Moving this responsibility into Qdrant changed the shape of the system.

The application now focuses on orchestration: parsing code, generating embeddings, storing metadata, and asking semantic questions. Qdrant handles vector indexing, similarity search, scoring, and retrieval.

That separation matters.

It made the pipeline faster, but also cleaner:
- fewer memory-heavy operations in the app layer
- clearer ownership between application logic and retrieval infrastructure
- a much better foundation for future RAG-style workflows (MCP Server?)

It doesn't have to be “always start with the most advanced tool.” Start with the simple architecture that lets you understand the problem. Then, when the system shows you where the boundary is, move the responsibility to the right layer.

In this case, vector similarity search clearly belonged in a vector database.

Qdrant turned out to be the right fit.

Repo: https://github.com/wbrycki/code-genius

Qdrant: https://qdrant.tech/

Monday, May 4, 2026

Do You Really Need Horizontal Scaling?

In times of hype around distributed systems, it’s tempting to scale right away. Using Spark, Flink, or Dask as a processing engine is often straightforward. 

But should we design for clusters from the beginning, or start simpler and evolve the system over time?

To effectively scale, we need to understand the data volume. Because horizontal scaling can often be an overkill.

Scaling is not free. The more distributed a system becomes, the higher its complexity. More components means more failure points, operational cost rises.

Vertical scaling is underrated in my opinion.

Newest machines have dozens or even hundreds of cores. Same with RAM.

The “just add more nodes” approach is often the default assumption. But what does adding more nodes really mean? Network overhead, serialization, harder debugging.

Image: Spark Data Scaling Horizontal and Vertical 

 

Example of ETL batch processing:


Spark runs in two main execution modes: 

  • local mode (driver and execution on a single machine)
  • cluster mode (distributed executors across nodes)

 

Reality Check on Horizontal Scaling

In distributed systems, moving data between partitions introduces a cost known as shuffling.

Shuffling is required to achieve proper data redistribution and balanced workload across partitions. It also occurs in single-node systems (Spark still operates with a logical distributed execution model, even when running locally). However, in cluster environments, shuffle becomes significantly more expensive due to network communication and coordination overhead.

Technical cost

  • network latency, data shuffling
  • In many cases, Spark performance is dominated more by shuffle behavior and data layout than by raw compute.
  • serialization/deserialization (can also happen on single node, for example local disk spillover)

Operational cost

  • Kubernetes, cluster management
  • monitoring, observability
  • deployments

Cognitive cost

  • debugging distributed systems
  • tracing
  • consistency issues

However, a single large machine is not a universal solution. Now we come to trade-offs.

With vertical scaling we have simplicity, lower latency, easier debugging.

However, vertical scaling is ultimately constrained by hardware limits.

At a certain point, the data size or workload characteristics simply exceed what a single machine can handle. The question is: are you actually at that point yet?

As Microsoft research point out: 

"that the majority of analytics
jobs do not process huge data sets. For example, at least
two analytics production clusters (at Microsoft and Ya-
hoo) have median job input sizes under 14 GB"

or 

that the majority of real-
world analytic jobs process less than 100 GB of input,
but popular infrastructures such as Hadoop/MapReduce
were originally designed for petascale processing."

from: Scale-up vs Scale-out for Hadoop: Time to rethink?

This research paper is from 2013 and now we are dealing with a lot more data, but not always and not everywhere - I still experienced under 100GB workloads in current times.


Now let's move to Horizontal scaling: we get scalability, fault tolerance (new Executor can re-start a task)

For some systems it's the only possibility, and we have to count the added complexity and operational overhead in. 

Making decision:

Are we CPU-bound?  -> scale vertically -> tune jobs -> then scale horizontally if not helping

Tuning jobs: data organization (e.g., Iceberg partitioning - when using Iceberg) often has higher impact than scaling compute. 

Are we really in Petabyte scale? -> scale horizontally, as single machine will probably not be effective

SLA (Service Level Agreement) Low Latency or High Availability requirements -> scale horizontally.

Does the operational cost align with business expectations? Maybe longer running job over night will also be tolerable. Cost per Job?

I would rather not scale in early stage system or up to couple TB of data. 
I would rather scale for massive datasets, HA requirements, or when dealing with streaming or real-time.

Even if we need to think about scaling from the beginning, it is still better to start simple and scale later. But always measure first.
It is useful to understand tools like autoscalers and Kubernetes taints & tolerations (in environments such as EC2 nodes on EKS), but it is equally important to know when not to use them.

As always, it is not only CPU and memory that matter, but also I/O bottleneck. A larger VM does not always translate into linear performance gains. And scaling is a decision, not a default.


Tuesday, October 29, 2024

Evaluation of role playing in LLM Systems based on chatGPT


Many of us already have had or will have an interaction with some form of Large Language Model (LLM), regardless of whether we want it or not. 

Direct or indirect. 

It could be a chatbot, a mail that we just got from someone, a recipe, detailed guidance on technical problems related to programming or general life advice.

When we ask a question, ChatGPT appears to us as a "wise machine" which can, in theory, help with everything. Of course help (advise, summary, explanations, guides, suggestions...) comes not in physical form, but rather text, image, video in some cases. which does not make it any more or less important.

 

AI System that generate dialogues, a chatbot. Is that what it is? Well, I think there is more.

 

 It is not only the system prepared for giving us the answers, it was instructed to behave in this on another way, to provide only high quality answers. 

 

Have you noticed chatGPT almost never answers in a rude way?

 

Even if I have a bad day and don't write "please" at the end :) And I've experienced already something else with different, not so well refined models. That would mean, for sure chatGPT was instructed, per default, to give us the nice, quality answers. To make it a better product. To act as an omniscient being, that likes to share the knowledge.

But that means it's already playing a role, because it's in that just described role. That would mean, we can ask ChatGPT to alter the role a little bit, or even preserve the "defaults" AND be at the same time, a debate opponent, negotiator, character in a story, investigator, or a diplomat. 

So now, ChatGPT having all the skills from before, can have also specific stance or personality.

Examples? How will I use it? Coming right up. 

Inspired by various posts about what I can use ChatGPT for, I came up with an idea. Let’s play roles: ChatGPT will be the interviewer for my current position, and I need to pass the interview. I’ll formulate it like this:

“Act as an interviewer for a Senior Engineer role and ask me questions. Then tell me if I got the job or not.”

Tuesday, September 24, 2024

Consider using Polars whenever you can

When working with large datasets, tools like Spark, Iceberg, and Trino have become popular in many data engineering workflows. They are robust, scalable, and capable of handling complex operations across distributed systems. However, there are cases where these heavyweight solutions might not be the best fit, especially when working in smaller environments or standalone systems. Enter pola.rs, a powerful alternative that offers speed and efficiency, particularly in certain scenarios.



Traditional architectures and the Bottlenecks
Let’s consider an architecture like Spark working in standalone mode only on one machine. Suppose we have data in a data lake (perhaps a CSV or Parquet file from another system) We load it into Spark to clean and process the data before analysis. Spark, though extremely powerful, comes with some trade-offs, especially in a standalone setup.
In that mode, Spark is limited to a single machine, no matter how powerful that machine is. This limits its scalability and overall efficiency when working with large datasets. Furthermore, the initial setup and resource overhead of Spark can be overkill for simple data transformations or cleaning tasks. That’s where pola.rs shines.

Why pola.rs is a Better Fit for this scenario
pola.rs, the Rust-based DataFrame library, is designed for speed and ease of use. When compared to Spark, especially in standalone environments, pola.rs can outperform Spark significantly, particularly for tasks that don't require distributed processing. Imagine you're loading a large CSV or parquet file with millions of rows to a data lake. 


Instead of starting a Spark cluster, which consumes a considerable amount of memory and time just to get going, you could use Polars to process the data much faster on a single machine.
Polars uses Rust’s zero-cost abstractions (code that is both expressive and efficient) and efficient memory management, allowing it to process data faster with fewer resources. For example, loading and cleaning data from a CSV or Parquet file can be done in a fraction of the time that Spark would take, especially for mid-sized datasets.

 

Benchmark performed against TPC-H, open source [img source: polars.rs]

Streaming with Polars
Another key area where Polars shows its abilities is working with streaming data. If your workflow involves importing data into a system like Microsoft SQL Server, pola.rs offers efficient ways to handle this in a streaming fashion. By leveraging Polars' streaming APIs, you can import large datasets in chunks, avoiding the need to load everything into memory at once (of course additional tool like efficient bcp utility will be additionally used, but combining it with Polars is possible!).

In addition, Polars can also be used for ETL workflows, efficiently processing data and transforming it as it's being streamed from the source to the destination. However we have to have in mind, that it’s not streaming-first library like Kafka or Flink.


When to use pola.rs?

When should you consider using pola.rs over heavier solutions like Spark?

  1. Standalone Environments: If you’re working with Spark Standalone, you might find that Polars is faster and more efficient for cleaning and preprocessing your data. In scenarios where the scale doesn’t justify the complexity of a Spark cluster, Polars is a great fit.
  2.  Mid-sized Datasets: Polars excels when working with datasets that fit comfortably on a single machine but are too large for in-memory tools like Pandas.
  3.  Streaming Workflows: If your data ingestion involves streaming into systems like Microsoft SQL or handling real-time data, Polars can be an optimal choice for managing this efficiently.


Downsides to Consider
While Polars is fantastic for many use cases, it’s not without some limitations:

  1. Lack of Distributed Processing: Unlike Spark, which can scale across many machines in a cluster, Polars is designed for single-machine workloads. If your data is truly massive and requires distributed processing, Spark or a similar framework may still be the better choice.
  2.  Ecosystem: Polars, while powerful, does not yet have the rich ecosystem of plugins and integrations that Spark has. If your project relies on many external libraries or data sources, you may end up with writing more custom code with Polars.
  3.  Big data table formats: no native support of formats like Apache Iceberg, Paimon etc. Need of conversion first. But, since it can fit between parquet file and writing into an iceberg table, in those scenarios this will be no downside (we still need to convert either way, can use streaming/batching approach and use Spark with code for Iceberg, or PyIceberg).


For many data engineering tasks, especially those involving standalone machines Polars offers a faster, more lightweight alternative to Spark and other traditional big data frameworks. It reduces overhead, offers excellent performance, and is ideal for working with moderately sized datasets. If you haven't yet explored Polars, now might be the time to give it a try and see if it fits your next project better than the heavier options in your current stack.

Wednesday, September 4, 2024

How (NOT) to upgrade Spark

Upgrade to Spark 3.3 finished already couple of months ago. Everyone was happy because of new features and performance improvements. But as time goes, new vulnerabilities are raised through the scanner. Oh, wait a minute, right... those are Spark dependencies. Time to upgrade Spark. Soon.



 There are couple of options. There is Spark 3.4 in stable version with features that some teams are waiting for. There was also Spark 3.5 in alpha version (for the time being). Easy choice. Going with version 3.4.

 Is there a spark-iceberg-runtime ready yet? Uff, yes, there is let's use it. Couple adjustments here and there, upgrade, couple other components, resolve versioning conflict and "BAM!", we have it. At least I think we have. Let's run those e2e tests.



Oh no... couple errors. Not big a deal. TimestampType without timezones is now TimestampNTZType, some fixtures changed, need to adjust. More or less. Couple different other adjustments, lets run again, green this time.

 Yes we have it. Ready to deploy, ready to celebrate. But did someone test it on prod yet? Ok, maybe just a main feature. All green...

...wait a sec... no it's not, there is some OOM, but how can it be, with 3.3 it was working?

Since upgrade start, a new Spark version came out - 3.5.1. That should be stable enough... Trying that out. Maybe it will bring also some interesting new features. Again, adjusting the code, building, same problem. Only that one point in code - maybe we can optimize that somehow so it doesn't throw OOM? Probably we can...


Many companies, who implement some sort pipelines with data processing by themselves have to deal with that problem. It's a trade-off. New versions are bringing new features, but - can be more fragile to another constellation of parameters, which were working correctly in previous versions. 

Another point is the proper use of Spark. Avoid unnecessary checkpoint(s), cache, persist here and there. Be aware of limits of the query plan length, do thing in another way, if possible.


Upgrading Spark is always a journey, full of both challenges and rewards. We've seen firsthand how new versions can bring exciting features and performance boosts, but also the occasional hurdle, like the unexpected OOM errors. However, with each upgrade, Spark continues to evolve, improving stability and expanding its capabilities.

Let’s approach this upgrade with the care it deserves—starting with thorough testing on smaller pipelines to ensure a smooth transition.

On Aug 10, 2024 a new version of Spark was released, 3.5.2. Improvements in Parquet timestamp handling, preventing OOM in some cases, as well as many other improvements. Making it a worthy consideration. Spark is constantly improving.

By staying in the same Spark version we loose the edge, we do not evolve.

By staying current, we’re not only solving today’s problems but also positioning ourselves for success as Spark continues to grow. Spark 4 is just around the corner, and being ready for it means embracing these updates now, leveraging the collective experience of the community, and preparing our systems for the future.

Let's move forward with confidence, knowing that each step brings us closer to a more robust, efficient, and future-proof platform.



Saturday, August 24, 2024

Data Mesh - key features

Data Mesh

Imagine a world without all the complicated data pipelines, without transformations, and without moving data elsewhere, yet still getting value from data. This is what Data Mesh is about, though there is, of course, more to it.

Data Mesh is not something you can buy. It’s a set of principles that we follow.

Data Mesh represents a decentralized approach to accessing data at scale and draws heavily from Domain-Driven Design.

It’s about deriving value directly from the source, rather than having data flow through processes to be valued later (with the added benefit of linking different models of those domains).

This domain orientation is somewhat like microservices in system architecture, but applied to data architecture.

Many current data architectures are built around ETL processes. ETL processes with fragile pipelines within the data lake.

Domains currently generate data for their own use, each with its own business processes, and don’t necessarily consider how analysts will extract value from that data later.

 

Data Lake teams are responsible for every part of the process, from data import, cleaning, and transformation to other representations, data pipelines in the data lake, and how the data appears in analytics and reporting.

These teams often become bottlenecks in the functioning of the data lake. We can’t just expand them indefinitely because it simply doesn’t work.

Do we even know what we’re discussing in daily meetings with so many topics and developers involved? OK, let’s split the team. But then… we have an even more complex structure that needs to communicate with itself.

 

Data Mesh involves both organizational transition and technology.

There is no specific technology that must be used. However, some technology is still necessary — meaning whatever fulfills our needs.

You could say you're implementing Data Mesh — or more accurately, applying the principles of Data Mesh, provided you have the appropriate technology and use it correctly.

We don't solve data management issues merely by purchasing a technical solution. It's more about how we use it, how people and processes interact with the data, as this is how we can address data quality and ownership problems.

By following Data Mesh principles, the issue of bottlenecks with highly specialized teams is eliminated.

Domain owners must take real ownership of the data and ensure it is usable for others who need it.

Is data from the entire domain accessible to everyone?

Or are only certain parts of it accessible (to a limited number of users)?

The technology should allow for unified (connectable) models with other domain owners. A model may connect only some of the data products from within the domains.


Data Mesh is not for everyone.

It addresses problems faced by larger organizations and the bottlenecks associated with Centralized Data.

Data Mesh is generally not intended for organizations with a small number of domains and/or data, it's designed for complex environments.

If you have a well-functioning data lake with just a few products and a relatively small amount of data, you might be perfectly fine without adopting Data Mesh principles.

As complexity grows, with more data sources and larger scale, the potential of Data Mesh increases.

 

Data Mesh is not implemented all at once.

Start small and improve. Build iteratively. Be use-case driven.

 Create Data Products. They should be relatable and linkable to each other. Later, you can hand them over to domain owners. Even if they are not currently discoverable and accessible by other domains, they might be in the future once the platform and other domains are ready for it.

Avoid large, sudden changes. Prevent user mistrust ("I will continue using the current solution because it works, and the new one is too complicated").

There is no magic button for "applying Data Mesh" — click and it’s done. This is due to technological challenges as well as the challenge of changing company culture.

 

The principles of Data Mesh complement each other.

The first and most important principle is Domain Ownership.

The other principles are essentially implications of this first one. Domains alone do not facilitate the registration of data products, access to them, or defining access policies. This is why we have the following three principles:

Data as a Product

When creating domains, we must somehow make them accessible to the outside world, treating them as products that can be used by others.

Products can be registered on the platform, discovered later, and used, potentially combined with other products to extract meaningful data.

… and these products must be registered and maintained on:

Self-Serve Platform

Which is essentially the technical execution of Data Mesh.

 Federated Computational Governance

  • Federated means that we have a federated group of experts from each domain who discuss and decide together on: how data is accessible, what exactly is accessible, and for whom, for which Data Mesh users/groups.
  • Computational means we can compute results based on different domain models.
  • Governance means how the policies are defined, whether we need to mask certain columns because some data should not leave the domain scope, and perhaps data cannot leave a certain region.

The biggest technical challenge is to create a truly self-service platform.

A self-service platform should not only allow for registering and finding products but also enable linking models (products) across domains. Such model connections can be realized through distributed query engines (like Trino or Spark SQL).

When implementing a platform, several questions and considerations arise:

  • How can we justify the increased cost of creating these domains?
  • Do we need new databases?

Perhaps not; it is still the same object storage we have now (we simply connect directly to the source).

  • Do we need to introduce new roles in the organization? What about the cost of that?

Not necessarily new roles; the right person needs to be empowered to use the self-service platform (it's about taking responsibility for managing the domain).

  • Significant investment in development

This may be true, but we are mainly using the same technologies as before, just in a different manner. We should also consider the cost of missed opportunities (if we do not adopt a Data Mesh architecture).

  • Production data directly accessible to others? That won’t work…

It doesn’t always have to be that way. Domains can provide the most relevant copy of the data, but it should still occur within the domain itself; similarly, if pipelines are needed, they remain within the domain.

Conclusion

For me, the key takeaway is embracing the culture of transformation that needs to occur. As a domain owner, people need to be responsible for the data to ensure it is of good quality, discoverable, and accessible.

You can have a great tool, but without transformation in the company and a shift in mindsets, it will be practically useless.

 

Check out also this webinar:
https://www.thoughtworks.com/about-us/events/webinars/core-principles-of-data-mesh/data-mesh-and-domain-ownership