The Library MMXXVI
Open shelves

A room for the thinking.

Everything in this room is given away. The paper that named the era, annotated for organisations. The long history of attention. The method in full. The books we keep returning to. Take your time, and take what is useful.

Frontispiece A timeline of attention

Two histories, converging.

Attention has been studied seriously for one hundred and thirty five years. The field of artificial intelligence discovered it formally in 2017. The two lineages converge on the same question and answer it from opposite ends. Our work sits at the join.

THE COGNITIVE TRADITION philosophy, psychology, neuroscience William James Principles of Psychology 1890 Donald Broadbent filter theory of attention 1958 Daniel Kahneman Attention and Effort 1973 Iain McGilchrist The Master and His Emissary 2009 THE JOIN 17.06.2017 THE MACHINE TRADITION machine learning, agentic systems 2014 Bahdanau et al. attention in neural translation 2017 Vaswani et al. Attention Is All You Need Now the agentic era The same question asked from opposite ends. The same answer arriving in the same year.

One hundred and thirty five years of attention The cognitive tradition asked how the human mind directs focus.
The machine tradition asked how a system might learn to focus its own.
Both arrived at the same answer in 2017.

§ 01 The method

The architecture of attention, applied to organisations.

/01i.
Attention Audit where leadership focus goes
Most companies spend their attention on the wrong things. The first job is to see clearly where it goes today and where it should go tomorrow.
/02ii.
The Encoder understanding the organisation as it is
We absorb the operating reality, encode it into a working model, and represent it back to leadership in a form they can act on.
/03iii.
Multi-Head parallel perspectives, integrated
Specialists working in parallel from different angles, operational, financial, cultural, technical, market, synthesised into one coherent view.
/04iv.
The Decoder strategy generated from understanding
Where understanding becomes action. The operating model, the deployment plan, the things that need to ship.
/05v.
Residual Connections preserving what already works
We do not tear down. The capability you already have skips past the bottleneck of redesign and continues forward, intact.
/06vi.
Positional Encoding context, timing, and place
Generic advice fails because it is not positioned to a specific moment. Every recommendation is encoded to where you are.
/07vii.
Layer Normalisation stable ground for deployment
Agentic systems land on stable ground or they do not land at all. We normalise the operational layer so the deployment is real, not performative.
/08viii.
The Transformer Block integrated, repeatable, alive
When all of the above are working together, the organisation has a complete unit that can repeatedly transform inputs into outputs. That is the work.
§ 02 The annotated paper

Reading Attention Is All You Need,
through the lens of organisations.

The paper that opened the AI era was a research artefact. We have spent considerable time reading it as something else: a guide to how attention, hierarchy, and parallel processing might be redesigned in the organisations of the next decade. What follows is a small selection of passages from the original paper, with our commentary.

From the Abstract "The dominant sequence transduction models are based on complex recurrent or convolutional neural networks. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely."
Our reading

The first sentence of the paper is a diagnosis. The dominant approach is too complex. The second sentence is a thesis. Simpler architecture, focused on the one thing that mattered (attention), would outperform. Almost every successful organisational redesign we have ever seen follows this structure. Diagnose where the complexity is unnecessary, identify the one mechanism that matters, build around that. Most companies inherit elaborate machinery from a previous era and never ask whether the machinery is still doing the work.

Section 3.2.1 "An attention function can be described as mapping a query and a set of key-value pairs to an output, where the query, keys, values, and output are all vectors. The output is computed as a weighted sum of the values, where the weight assigned to each value is computed by a compatibility function of the query with the corresponding key."
Our reading

This is the technical heart of the paper, and it has a remarkable organisational analogue. An organisation is also a function that maps queries to outputs by weighting which information sources are most compatible with the question being asked. A good organisation routes the right query to the right person with the right context, and weights their input correctly. A bad organisation routes everything to everyone and weights nothing. The transformer's contribution was to make this routing explicit, parallel, and differentiable. The same operations could redesign how decisions get made inside companies.

Section 3.2.2 · Multi-Head Attention "Instead of performing a single attention function with d_model-dimensional keys, values and queries, we found it beneficial to linearly project the queries, keys and values h times with different, learned linear projections."
Our reading

Multi-head attention is the discovery that one perspective is not enough. The model gets better when several different attention mechanisms run in parallel, each looking at the same data through a different learned lens, and the results are concatenated. This is exactly how good consulting works, when it works. A senior team brings several perspectives to the same problem in parallel, integrates them, and produces a synthesis that no single perspective could have reached. Most organisations operate as single-head systems. They have one dominant lens, usually financial, and miss everything the other lenses would have caught.

Section 3.4 · Why Attention "A single attention layer connects all positions with a constant number of sequentially executed operations, whereas a recurrent layer requires O(n) sequential operations. Self-attention layers are faster than recurrent layers when the sequence length is smaller than the representation dimensionality."
Our reading

Recurrent networks process information one step at a time, like a hierarchy passing memos up and down. Self-attention sees the whole context at once. This is the deepest structural insight of the paper for organisational work. Most companies are built on recurrent communication. Information goes up the hierarchy, gets compressed, comes back down, gets misinterpreted. The companies that will thrive in the agentic era will be the ones that flatten this. Where context is shared, where everyone sees the whole sequence, where decisions can be made closer to where the work happens.

§ 03 The reading list

What we read, and why we read it.

A short list, updated quarterly. Not a syllabus on artificial intelligence. A syllabus on people, work, attention, and the longer arc of how technology either lifts or diminishes us.

2009 Iain McGilchrist
The Master and His Emissary

Two ways the brain attends to the world. The mode you choose determines what you can see. Foundational text on why attention is the substrate of every other capacity.

1958 Hannah Arendt
The Human Condition

The distinction between labour, work, and action. Reads as if it were written for the agentic era. The most important book about what humans are for that we know of.

2002 Carlota Perez
Technological Revolutions and Financial Capital

The pattern that connects every technology revolution to the social one that follows. Helps locate where we are in the AI cycle.

1973 Peter Drucker
Management: Tasks, Responsibilities, Practices

The original text on knowledge work and what it asks of organisations. Drucker saw most of this fifty years ago and most companies still have not caught up.

2017 Vaswani et al.
Attention Is All You Need

The paper the firm is named for. Eight authors, eight pages, the architecture that opened the modern era of artificial intelligence. We annotate it above.

1972 Stafford Beer
Brain of the Firm

The cybernetic view of organisations. How information flows, how feedback loops work, how a company is more like a living system than a machine.

1999 Stewart Brand
The Clock of the Long Now

Thinking in centuries rather than quarters. The intellectual posture every senior leader needs and almost none cultivate.

2018 James Bridle
New Dark Age

The cost of mistaking computation for understanding. A counterweight to the techno-optimism that dominates AI discourse.

1964 Marshall McLuhan
Understanding Media

The medium is the message. Every new technology rewrites the human relationships around it. AI is the most consequential medium since print, and we are not yet asking the right questions about it.

2023 Karen Hao
Empire of AI

The political economy of the AI industry, told without flattery. Required reading for anyone deploying these systems into their organisation.

The way back

The house is through here.

The studio, the offer, the writing, and the door are on the front page. If something on these shelves was useful, that is the whole point.