Skip to main content

Data and Algorithms

Introduction

Data is used by all layers of the technology stack. There are different types:

  • Metadata — data that describes other data: the time a message was sent, who sent it, from where
  • Personal data — data that can be linked to a specific person
  • Open data — data shared openly under defined conditions so anyone can use it

A lot of data contains information about behaviour: where people go, what they find interesting, who they interact with, whether they are healthy or sick. Algorithms recognise patterns in all of this, and that pattern recognition is what powers analysis and decision-making at scale. Data and algorithms are what make services work — but they can also be misused for manipulation, automated decision-making, and influence at scale.

Trusts, cooperatives, and commons

Data is a resource, and it's unevenly spread. Corporations sit on huge stashes of it, accumulated by default from every interaction with their platforms; movements and unions barely have any, because nobody was collecting it on their behalf. Data trusts, cooperatives, and commons exist to correct that imbalance — holding data collectively, governed by the people it's about, rather than by whoever happened to build the platform that captured it.

What you choose to collect decides what you can actually see. An organisation that collects data about its members is monitoring humans. One that collects data about capital, power, and externalities instead — who owns what, who's causing harm, what it actually costs — is monitoring the things worth watching. The same resource, pointed in a different direction, becomes a tool for accountability instead of a liability waiting to leak.

In practice

The public stack means data stays on infrastructure you control. No third party processes it, no algorithm profiles your members, and no vendor can change the terms under which data is stored or used.

Data interests — what we choose to look at, and why. The data interests of a public stack — especially a self-organising, non-profit one — are fundamentally different from a company's.

  • No member data — watching our own members doesn't make sense at all: we don't collect metadata beyond what's needed to function, and we don't do algorithmic profiling. There's no need to capture more attention, and no need to see behaviour coming — we hope our members stay stubbornly unpredictable
  • Organisational data — there's no revenue here, so tracking it the way a business would makes no sense. We design our own metrics instead: the ones we actually wish to be measured by, like social impact or ecological impact — the externalities a for-profit would never bother to price in
  • Domain data — what we do need to watch more is exploitation: how it moves through buildings, organisations, and systems, and how it gets acted upon. To build power, we need to know what's actually happening: how quantitative and qualitative changes drive that exploitation, not just that it exists

Data minimisation — not collecting something is the simplest way to protect it, and the cheapest way to store it. Retention is kept to a minimum, for security reasons as much as environmental ones: what isn't stored can't be handed over, leaked, or misused later, and it isn't sitting on a disk somewhere burning energy for no reason.

Open formats — data can be exported, migrated, and used independently of the software that created it. This is non-negotiable: a tool that doesn't support it isn't evaluated further, no matter what else it offers.