# Project Description
In collaboration with engineers from the Zurich University of Applied Sciences (ZHAW), I am conducting systematic research into what machine vision means in practice: its limitations, advances, recurring patterns and those peculiarities that only become apparent through repeated, controlled observation. The subject of the investigation is not the image as such, but the gap between image and description – the gap that opens up when a language model attempts to translate visual reality into language.
The methodological approach is deliberately kept simple: a large language model is tasked with describing randomly selected live cam images in text. The randomness of the image selection is intended to prevent curated or predictable motifs from distorting the result and instead to generate a range of situations in which the strengths and weaknesses of the model are equally apparent: the familiar and the unexpected, the unambiguous and the ambiguous, the describable and that for which language seems to fail.
The newsticker visible in the stream above is renewed daily, forming a kind of machinic diary of seeing — an ongoing record of what the model perceives in the world, or more precisely: what it deems worth perceiving and renders into words.
Clicking on any of the scrolling texts opens the described image alongside the generated text, inviting direct comparison: Is what is described accurate? What is missing? What has been invented? And what does the discrepancy reveal about the model's inner logic?
When an artificial intelligence describes a photograph, something epistemically peculiar comes into being: not the report of an eye, not the testimony of an experience, but the output of a statistical pattern-recognition process cast into linguistic form. What emerges in this process deserves closer examination — both with regard to what succeeds, and to what remains structurally foreclosed.
## What the AI Sees
Multimodal language models do not process images through perception in any phenomenological sense, but rather by transforming pixel matrices into high-dimensional vectors. These are correlated with linguistic representations distilled from vast text corpora. The model sees in the sense that it recognizes visual features — edges, colors, textures, spatial relations — and associates them with learned concepts. In doing so, it produces descriptions that are statistically plausible: what is typically said about images of this kind.
This gives rise to remarkable capabilities. Objects, persons, scenes, moods — all of these can be named with considerable precision. Cultural contexts are recognized, pictorial genres identified, compositional choices described.
## The Structural Limits
Yet this is precisely where the limits begin. An AI's description is fundamentally referential: it points to what an image depicts, rarely to what it is or does. The difference between a likeness of a face and the portrait of a person — with all its intimacy, its historical moment, its ethical dimension — remains largely inaccessible to a model that lacks embodiment and temporality.
To this must be added the problem of context. A photograph by Dorothea Lange may be technically described with accuracy — an exhausted woman with children, a migrant camp — yet the semantic surplus of the image, its documentary force, its place in collective memory, can only be unlocked through cultural and historical knowledge that is not itself encoded in visual patterns.
Particularly instructive is the question of the non-depicted. Photographs also signify through omission, through the frame that determines what is left outside. What an image conceals or withholds is structurally inaccessible to machinic description — because it carries no visual features that could be recognized.
## Description as Translation
The process can be understood as a form of translation — from one sign system (the visual) into another (the linguistic). And what holds for every translation holds here too: there is fidelity, and there is loss. What the AI delivers is a plausible, occasionally illuminating, never complete paraphrase of what an image shows. It is neither seeing nor reading, but something third: a statistically generated equivalent that, at best, can serve as a point of departure for human interpretation.
The AI's seeing is therefore not seeing in the sense of awareness, but a highly sophisticated recognition — without wonder, without being affected, without what Roland Barthes called the punctum: that point in an image which pricks, which strikes, which wounds. Whether such categories will ever become machinable remains an open question — and perhaps the most interesting one this field has to offer.