Selected work04 of 04
Filmstock Responsiveness
A controlled study of how image models respond to filmstock, lighting and scan language in prompts, written up as a 31-page paper and replicated on a newer model.
- Role
- Sole author — design, analysis, paper
- Year
- 2026
- Status
- Research
- Stack
- Python, PCA, MANOVA, PERMANOVA, ComfyUI

Overview
Prompts for photographic images lean on the vocabulary of film: stocks, lighting setups, development and scanning. This study asks whether image models respond to that language, and by how much.
The first study compared ChatGPT Image 2 with Grok Imagine on a full-factorial grid — 288 normalised images, 195 extracted image features and 432 blocked prompt-response contrasts — analysed with PCA, MANOVA and PERMANOVA.
A September replication compared Images 2.0 with Images 2.5. Prompt-response distances were 24.0% lower for filmstock, 14.0% lower for lighting and 36.9% lower for development and scan wording (Holm-adjusted p < .001), and repeats of the same prompt were 41.5% closer together. The study measures these vocabularies only; it does not establish a general decline.
What I did
I designed the experiment, generated and normalised the image set, built the feature-extraction and statistics pipeline, and wrote the paper. The code is MIT-licensed; the data and writing are CC BY 4.0.
Numbers
- normalised images in the first studyfilmstock-responsiveness-benchmark, README
- 288
- image features extracted per imagefilmstock-responsiveness-benchmark, README
- 195
- pages in the replication paperstudies/2026-09-images-2-5
- 31
Figures



Related research
Criticality in sparse coding — a simulation study of phase transitions in a binary sparse-coding model, sweeping sparsity and overcompleteness across 600 Gibbs-sampling runs.
Heart-rate variability in concussion baselines — as a research assistant at Mount Sinai’s Abilities Research Center, I co-authored an abstract presented at the ACSM Annual Meeting in 2025 and published in Medicine & Science in Sports & Exercise.
