Show Me Your Dataset and I Shall Tell You Who You Are

2025

Publication History:

“Show Me Your Dataset and I Shall Tell You Who You Are,” In Stainslas Chaillou, Artificial Intelligence and Architecture: From Research to Practice,  Second Edition, 194-199. Basel: Birkhäuser, 2025

The text posted here is a preprint draft and is significantly different from the published version. Please only cite from copy in print

As we all know, since the summer of 2022 everyone has been obsessed with generative AI. First, because with few exceptions, nobody had seen it coming — I didn’t, for one; so, when it came it was a shock. Second, because some find it promising, and some frightening, never mind that for the time being what generative AI can do amounts to, practically, almost nothing. But before we try to come to terms with what generative AI can do for, or against, the design professions, we must clear up what kind of artificial intelligence we are talking about — as there are many artificial intelligences, often confused under the same name.

A branch of computer science called “artificial intelligence” has been around since 1956. In 1960 Marvin Minsky famously defined artificial intelligence (at the time, not capitalized) as a “general problem-solving machine”; a definition that held good until very recently, and possibly still does. Yet since those distant times, the discipline has already risen and fallen, and reinvented itself many times. Artificial intelligence was very popular in the 1960s, including among architects (see, for example, the groundbreaking book by Nicholas Negroponte, “The Architecture Machine,” first published in 1970); this was the time when many thought that computers could do almost everything. As of the mid-1970s, however, it appeared that computers of that time could do little more than nothing; so both cybernetics (Norbert Wiener’s new science of “control and communication in the animal and the machine,” born in 1948 and partly overlapping with early computer science), and AI itself, were quickly forgotten and relegated to the dustbin of technical history — and of design history.

The digital turn that changed the history of architecture in the 1990s was driven by small computers used as drawing machines, not by big computers used as thinking machines. Computers in the ‘90s made drawings — including some very special kinds of drawings we could not have drawn without computers; but in the 1990s nobody even tried to use computers to solve design problems, and the precedents of Wiener, Minsky, and even Negroponte were more than forgotten — they were entirely removed from living memory, and they did not play any role in the rise of CAD-CAM, digitally intelligent design, spline modeling, BIM, or parametricism.

Things are changing now because artificial intelligence, — seen again, as in the 1960s, as a general problem-solving machine — is back, and this time around it actually kind of works. The recent revival of artificial intelligence came from a branch of computer science that had been tested in the 1960s, then abandoned in the 1970s, based on so-called neural networks; this mode of computation used to be known as “connectionist,” and today it is better known as machine learning. Building on that, the second coming of AI in the visual arts — after its first coming, in the 1960s — is due to a new kind of neural network, called generative adversarial networks, or GANs, which turned out to be very good at processing images.

GANs, and StyleGANs in particular, have been known since 2014–2016 and have been used by artists and designers since at least 2018. However, the recent hullabaloo over generative AI was due to the release of new, user-friendly text-to-image applications, such as Midjourney , and DALL·E, by OpenAI, in the spring and summer of 2022. And of course, everyone has been mesmerized by generative AI since the release of ChatGPT in the fall of 2022.

Old GANs, which artists and designers have been experimenting with for almost 10 years, and today’s generative AI, are based on a similar technical logic, and this is where it is capital for us, in the design professions, to make certain that we understand the spirit of the game — to find out how these technologies work, what they can do, and what we could try to do with them.

The technical logic of generative AI is by now well known. GAN artists must first gather a corpus of related images that are pertinent to their project; this corpus, called a “dataset,” is then fed into machine-learning algorithms that ferret out what these images have in common. The result of this inductive process is a mathematical matrix that computer scientists call a “latent space,” and that roughly corresponds to what philosophers would call the definition, idea, formula, or quintessence of the original dataset. So, for example, if we feed the system a dataset of images of dogs, the “latent space” will be the definition of an ideal dog. This definition, which today is neither verbal nor visual but mathematical, can in turn be applied to different tasks: first, analytic (showing the system a new picture, and asking: “Is this a dog?”); second, generative (asking the system to produce realistic images of dogs that do not exist, or fake dogs); third, combinatory (asking the system to merge two or more datasets in given proportions, producing, for example, pictures of a canine quadruped that looks 80% dog and 20% coyote).

 Now, there is an interesting variation of this combinatory mode, where generative algorithms are asked to apply some general aspects of a dataset not to a set but to a single individual — thus transforming this individual into a new one that still looks recognizably similar to the exemplar from which it derives, but at the same time shows some features of an external dataset, to which it is being assimilated. The general traits that define a visual set and are common to all its individuals, but pertain to none in particular, are akin to what western classical artists, starting in the Renaissance, used to call a “manner” or a “style”; not surprisingly, the computer scientists who first developed these processes of transformation and assimilation between an individual and a set decided to call this operation a “style transfer.” The term has stuck and is now in current use.

 The wave of generative AI that has swept the planet since the spring of 2022 has not significantly altered this conceptual framework, — except that the new chatbots, meant for the general public, are not based on specific, purpose-built datasets, but on huge and generic archives of labeled images (text-and-image pairs) scraped from the internet, called large language models. Evidently, due to their provenance, large language models can only reflect what the internet has already gathered. Cross-modal image generators based on large language models function by the retrieval, aggregation, and averaging of precedent; and, as the precedent from which they derive is a crowdsourced collection of images, generative AI is in this instance a machine that automates the imitation of visual commonplaces.

For example: if I type the prompt “Switzerland” into the latest version of ChatGPT, what I get is in fact an automated, visual opinion poll, showing what Switzerland looks like in the mind of most people (Fig. 1). This may be of some interest to sociologists, or help a tourist office to design a new website, but that would appear, prima facie, of very limited practical interest to the design professions. This is the wisdom of crowds, or the opinion of most, translated into one synthetic picture that is in turn dependent in full, and exclusively, on the dataset from which it derives.

Conversely, what is of interest to designers, I would suggest, is the potential of a machine that has long last acquired a capacity to successfully automate imitation. Regardless of the way the original dataset has been put together, generative AI can produce new images that are similar to, but different from, all the images in the original dataset. Generative AI imitates the dataset it was trained on — meaning, the datasets we fed into it. Until recently, creative imitation was seen as a privilege of the human mind. No longer. This is truly a major technical breakthrough, and one with vast epistemological, philosophical, and scientific implications.

But what’s the use of it in design? Do today’s designers need a machine to imitate a model, a set of models, or some abstract features common to an entire set of models — for example, to imitate an artistic style? Which designers would need someone else’s intelligence, never mind if artificial, to imitate someone else’s style — or even their own style? Yet, oddly, many well-known design offices around the world are doing just that, right now. And there may be a logic in it. Because this is what AI does; this is the way artificial intelligence works, and this should give us pause, and make us think.

Generative AI does not create images out of thin air. Instead, it imitates a carefully curated dataset of images we have fed into the machine, and generates new images based on some criteria or parameters of imitation we have set. Is this not what human intelligence always did and still does? Artificial intelligence is here to remind us that there is no creation without imitation, no invention without convention, no design without a tradition, no expression without a language, no project without sources. Artificial intelligence is here to remind us that the first step of each creative process is the awareness, then the acknowledgement, of precedent — of a precedent we relate to, never mind if we like it or not. Every dataset is a canon, every canon is a dataset, and every dataset is a collection of models that someone, at some point, must have chosen (Fig. 2). And regardless of the technology we use, every canon is by definition exclusionary: every time we define a canon — every time we build a dataset — we put something in, and we kick someone out. We may not like that, but that’s the way it is — and it always was, long before the invention of generative AI. And today, generative AI is here to remind us that the first step in every creative process is the invocation and selection of a dataset, and that there can be no creation without some precedent of reference — and without some reference to precedent. Precedent is that canon of our own choosing that gives meaning to our voice — for better or worse. Generative AI is here to remind us that we cannot escape precedent — but we can at least try to make it clear that we know there is one, even when we don’t like it.

Again, many would argue that many cognitive theories, and many art theories, Western and not Western, have always said the same — or at least acknowledged that. And that is true, but with a difference. Humans can be unaware of the sources of their inspiration; they often are. Humans can also forget and even deliberately hide, or obfuscate, the references — the histories, the traditions, the conventions — they have in mind. In short, humans can be dumb, or cheat. Generative AI is brutally transparent. Art created with AI always starts with a dataset. If generative AI works, it is because someone, somewhere, has made a dataset, and chosen a canon. There can be no anxiety of influence when we design with AI. Show me your dataset, and I shall tell you who you are.

As a design critic, I do not necessarily embrace this transparency. But I am glad that the rise of generative AI is forcing us to confront, once again, the role of precedent, tradition, and history in all creative and artistic expressions. This is, in my opinion, the main contribution of generative AI to contemporary design, and to design theory.

 

 

Publication

Artificial Intelligence and Architecture

Citation

Mario Carpo, “Show Me Your Dataset and I Shall Tell You Who You Are,” In Stainslas Chaillou, Artificial Intelligence and Architecture: From Research to Practice,  Second Edition, 194-199. Basel: Birkhäuser, 2025