# Project Description

In collaboration with engineers from the Zurich University of Applied Sciences (ZHAW), I am conducting systematic research into what machine vision means in practice: its limitations, advances, recurring patterns and those peculiarities that only become apparent through repeated, controlled observation. The subject of the investigation is not the image as such, but the gap between image and description – the gap that opens up when a language model attempts to translate visual reality into language.

The methodological approach is deliberately kept simple: a large language model is tasked with describing randomly selected live cam images in text. The randomness of the image selection is intended to prevent curated or predictable motifs from distorting the result and instead to generate a range of situations in which the strengths and weaknesses of the model are equally apparent: the familiar and the unexpected, the unambiguous and the ambiguous, the describable and that for which language seems to fail.

The newsticker visible in the stream above is renewed daily, forming a kind of machinic diary of seeing — an ongoing record of what the model perceives in the world, or more precisely: what it deems worth perceiving and renders into words.
Clicking on any of the scrolling texts opens the described image alongside the generated text, inviting direct comparison: Is what is described accurate? What is missing? What has been invented? And what does the discrepancy reveal about the model's inner logic?

01
automated image description (ai). Text below:An office space with a wooden desk and computer setup is shown in the center of the image. A whiteboard with notes and stickers is on the left wall, and a blackboard with white writing is positioned near the computer. A printer is situated on the desk to the right of the computer, and a small shelf above it holds a few items. A door is visible on the right wall, and a second desk with a computer and telephone is located beside it. The room has various electrical cords and wires, and a light fixture hangs from the ceiling. A toilet is visible in the background near the telephone. A dark-colored rug is placed near the doorway. The walls are painted white, and the overall impression is of a functional workspace. A small piece of clothing is draped over the left wall, and a metal shelving unit is partially visible on the left side of the image.



Machine Vision and Its Limits: On the Description of Photographs by Artificial Intelligence.

When an artificial intelligence describes a photograph, something epistemically peculiar comes into being: not the report of an eye, not the testimony of an experience, but the output of a statistical pattern-recognition process cast into linguistic form. What emerges in this process deserves closer examination — both with regard to what succeeds, and to what remains structurally foreclosed.

## What the AI Sees
Multimodal language models do not process images through perception in any phenomenological sense, but rather by transforming pixel matrices into high-dimensional vectors. These are correlated with linguistic representations distilled from vast text corpora. The model sees in the sense that it recognizes visual features — edges, colors, textures, spatial relations — and associates them with learned concepts. In doing so, it produces descriptions that are statistically plausible: what is typically said about images of this kind. This gives rise to remarkable capabilities. Objects, persons, scenes, moods — all of these can be named with considerable precision. Cultural contexts are recognized, pictorial genres identified, compositional choices described.

## The Structural Limits
Yet this is precisely where the limits begin. An AI's description is fundamentally referential: it points to what an image depicts, rarely to what it is or does. The difference between a likeness of a face and the portrait of a person — with all its intimacy, its historical moment, its ethical dimension — remains largely inaccessible to a model that lacks embodiment and temporality. To this must be added the problem of context. A photograph by Dorothea Lange may be technically described with accuracy — an exhausted woman with children, a migrant camp — yet the semantic surplus of the image, its documentary force, its place in collective memory, can only be unlocked through cultural and historical knowledge that is not itself encoded in visual patterns. Particularly instructive is the question of the non-depicted. Photographs also signify through omission, through the frame that determines what is left outside. What an image conceals or withholds is structurally inaccessible to machinic description — because it carries no visual features that could be recognized.

## Description as Translation
The process can be understood as a form of translation — from one sign system (the visual) into another (the linguistic). And what holds for every translation holds here too: there is fidelity, and there is loss. What the AI delivers is a plausible, occasionally illuminating, never complete paraphrase of what an image shows. It is neither seeing nor reading, but something third: a statistically generated equivalent that, at best, can serve as a point of departure for human interpretation. The AI's seeing is therefore not seeing in the sense of awareness, but a highly sophisticated recognition — without wonder, without being affected, without what Roland Barthes called the punctum: that point in an image which pricks, which strikes, which wounds. Whether such categories will ever become machinable remains an open question — and perhaps the most interesting one this field has to offer.

02
automated image description (ai). Text below:This black and white image captures a cluttered and possibly under-construction room. A black trash can dominates the left side of the image, with a partially visible white object near it. In the center, a white ladder leans against a pile of debris, including cardboard boxes and other discarded materials. On the right side, a metal frame stands tall, seemingly a part of an structure to be repaired or constructed. The walls of the room appear damaged or unfinished, and scattered debris covers the floor. A water bottle is visible near the ladder and a white plastic bucket is near the frame. The timestamp on the image indicates that the scene was captured between 2026 and 2023, during daytime.