https://openai.com/blog/multimodal-neurons/
ML Training on images and text together leads to certain neurons holding information of both images and text – multimodal neurons.
When the type of the detected object can be changed by tricking the model into recognizing a textual description instead of a visual description- that can be called a typographic attack.
Intriguing concepts indicating that a fluid crossover from text to images and back is almost here.