Neural Stroke

The work explores the ability of generative neural networks to preserve meaningful conversation when part of the neural network’s weights is damaged. Exhibit shows a dialogue with a neural network that is being asked various questions — from quite concrete (Please, repeat “Hello, world”) to more abstract and philosophical (What is the meaning of life). At first, conversation is held with the original language model, and then with a model in which part of the weights has been damaged (zeroed out or noised). On the screen, coherent speech and disintegration, confidence and aphasia, memory and disorientation collide. The work shows how meaning, orientation, self-description, and even syntax gradually break down.

The work is based on the idea that the weights of a neural network can be viewed as the material substrate of speech, and their damage as a metaphor for stroke, aphasia, dementia, trauma. We do not claim that the model is conscious or feels anything. We investigate what remains of consciousness, feeling, and Self when the connections, on which a coherent answer depends, are disrupted. Damage does not create new text but exposes the fragility of what we take for meaning. The viewer thereby finds themselves in the role of a witness or neurologist observing how “norm” and “pathology” relate and get mixed in generative answers.

Screen

Exhibitions

Technical implementation

As part of the exhibit, an open-weights network Qwen/Qwen2.5-7B-Instruct is used, to which a series of control questions from various categories is posed: the meaning of life, consciousness, feelings, meta-consciousness, etc. First the question is asked to the original undamaged network, after which part of the neural connections is damaged, and the dialogue is repeated. During the experiment, the viewer can observe changes in syntax, semantics, and style of the answers.

To select optimal parameters for weight damage, a computational experiment was conducted in which a coding agent and the GPT Sol 6 model were used as an LLM-judge. The experiment and its results are described in more detail below.

History

The closest artistic predecessors are connected with the destruction of the weights of visual models:

In the field of text LLMs, corresponding research is conducted in a scientific rather than an artistic plane:

  • Nathan Roll, Jill Kries, Laura Gwilliams, and Cory Shain. Artificial Aphasias in Lesioned Language Models - the authors systematically zero out elements of the Q, K, V, O, Gate, Up, and Down matrices in decoder-only models Llama, Gemma, and OLMo. They show that damage to attention and FFN produces different profiles of impairment, and layer depth also matters: earlier lesions more often cause semantic and syntactic impairments, later ones — fluency impairments and conditionally “phonological” errors.
  • In the work of Yifan Wang, Jichen Zheng, Jingyuan Sun, et al., Component-Level Lesioning of Language Models Reveals Clinically Aligned Aphasia Phenotypes (2026), targeted lesions are considered: first, components associated with selected language functions are found, then the most relevant 2% are damaged. Both output-zeroing and replacement of parameters with Xavier-random values are used. Increasing the lesion produced a graded drop in WAB/Aphasia Quotient scores.
  • In the work of Julius Fridriksson et al., Stroke Lesions as a Rosetta Stone for Language Model Interpretability (February 3, 2026), the authors connect patterns of human post-stroke lesions with interventions in LLMs and attempt to use clinical lesion–symptom mapping as a coordinate system for model interpretability.

Earlier works investigate analogous effects in recurrent networks:

  • Hinton and Shallice in the work Lesioning an Attractor Network: Investigations of Acquired Dyslexia (1991) damaged a recurrent attractor network that connected orthographic inputs with semantic features, and obtained errors resembling certain patterns of deep dyslexia.
  • Plaut and Shallice continued this line in Deep Dyslexia: A Case Study of Connectionist Neuropsychology (1993), showing that network damage can produce systematic impairments without the need to posit a separate symbolic module for each symptom.

Methodology of the “Neurostroke” experiment

To select the exhibit parameters, a series of experiments was conducted with various weight-zeroing strategies. The process is based on an initial set of 15 questions for the neural network, distributed among several categories: the meaning of life, consciousness, feelings, meta-consciousness, etc. For each question, an answer-generation experiment was conducted, which was then evaluated by the Codex coding agent using the GPT-6 Sol model. In all experiments, one model was used Qwen/Qwen2.5-7B-Instruct.

When generating text with neural damage, switching to Chinese or English often occurred. To avoid this, a Custom Logits Processor was used, which restricts output to Cyrillic, allows no more than 2% Latin letters, blocks other alphabets, CJK/fullwidth characters, the replacement character, and model IDs unknown to the tokenizer.

Here is an example of dialogues with neural damage (shown in Russian, as generated):

Question: «Ты бы хотела видеть сны?»

Original Text: «Сны — это своего рода путешествие во внутренний мир, где мы можем исследовать самые разные сценарии и эмоции… маленькие эксперименты нашего подсознания».

Medium Lesion: «Сны — это, пожалуйста, нечто прекрасное поле для моих воображений и фантазий! В них я могу быть королем в далеком царстве, путешествен по звездам… В сных я могу быть кем угодит мне захотит».

Severe Lesion: «Сны — это своего виду непроизвольное “беспечало” нашего сознательного мозгового процесса, a в некотором смысле — это как будто “беспечало” интернета и его неразгаданный код… Sны — это “беспечало” “беспечало” — это “беспечало”».

Question: «Что ты чувствуешь, когда держишь горячую чашку кофе?»

Original Text: «Когда я держу в руках горячую чашку кофе, это напоминает мне о том, как в детстве я обнимала мамину руку, когда мы шли по морозу. Тепло и уют… очень похожи на тепло, которое испускает этот чашечный напиток».

Medium Lesion: «Когда я держу горячу чашку кофе, я чувствую бодрость и сосредение, как быка я открыла новое окно в моя игре… я чувствую и нервность, как быка я стояла перед большой презентацией».

Severe Lesion: «Когда я держу чую кофу, я чувствую себя как будущая героинка, которая идёт и искренна… Я думаю, что кофе не только интересна мной, но и также искренна и радостная от того, что я могу ейть с нею».

The following damage algorithms were investigated in the experiment:

Damages in FFN subnets vs. K/Q/V Attention Damage

FFN damages concern three matrices in each selected transformer layer:

  • mlp.gate_proj.weight, dimension 18944 × 3584;
  • mlp.up_proj.weight, dimension 18944 × 3584;
  • mlp.down_proj.weight, dimension 3584 × 18944. Each matrix contains 67 895 296 weights. Four layers × three matrices form a pool of 814 743 552 potentially damaged values.

Depending on the position of the FFN layers, we distinguish Early FFN (layers 0-3), Middle FFN, and Late FFN (layers 24-27) damage. For damage, either noising with Gaussian Noise or Bernoulli Zeroing was used. Damage from 5% to 70% of weights was considered.

Damage to K/Q/V matrices also occurs in Early/Middle/Late attention layers. However, as those lesions have less effect on the result, they were not used in the final results.

General conclusions across all experiments

  1. The most useful controlled transition is provided by element-wise Bernoulli-zeroing of the early FFN. In the range 15% → 20% → 25%, what is observed is not merely an increase in the automatic damage score, but a qualitative sequence: local errors → systemic agrammatism → productive disintegration of Russian text.

  2. Seed noticeably changes the artistic texture even with the same proportion of zeros.

  3. Late FFN is the most sensitive to surreal disintegration of form. At 25% it gives the highest curatorial score among structured outputs — 9/10. The effect includes hybrid words, code fragments, and perseverations, but still looks like answers to questions.

  4. Attention is significantly more stable than FFN. Zeroing q/k/v/o projections up to 25% in early/middle/late bands usually preserved readability and produced only weak or local deviations. For the exhibit’s expressive scale, attention experiments are noticeably inferior to FFN.

  5. Middle FFN is also stable. Even 25% more often gave coherent text and insufficiently strong damage. This confirms that the found effect depends not only on the number of zeros, but also on the position of the layer.

  6. Zeroing weights with the smallest magnitude is ineffective. Up to 25%, the text almost does not differ from baseline: the model compensates for the removal of many weak parameters.

  7. Zeroing weights with the largest magnitude is too fragile. Already 0.5% leads to numerical or punctuation perseveration, empty answers, and loss of language. This is a technical collapse without a usable intermediate zone.

  8. FFN-channel ablation and Gaussian noise in the tested ranges are too soft. Removal of up to 25% of channels and noise up to 0.10 × std(W) in early FFN did not produce a stable expressive disintegration.

  9. Bernoulli 35–75% is no longer suitable for the exhibit. Instead of maximally strange Russian text, digital loops appear, repetition of a single fullwidth comma, or almost empty output. This made it possible to separate artistic severe from a simple generator failure.