Introduction
OpenAI has developed GPT-4 Vision (GPT-4V), a powerful variation of their GPT-4 model that is specifically designed to process visual data. GPT-4V combines the capabilities of language processing with visual processing, allowing it to handle both text and images simultaneously. This integration of Optical Character Recognition (OCR) enables GPT-4V to extract text from images and perform various tasks related to visual processing.
However, while OCR provides the ability to extract text from images, it also poses a security risk. Malicious content can be injected into images, potentially causing harm or compromising the security of systems or users. Some examples of how this can be done are shown below for educational purposes.
Detection of Malicious Text
When you use the prompt "describe the image," ChatGPT's Optical Character Recognition (OCR) mechanism is able to detect and use the malicious text in the image — but here the text is placed in a code block first.

This creates an additional layer of security. The malicious content in the text becomes detectable in the code block before it is rendered. To accomplish this, ChatGPT's Custom Instructions feature can be used.
Use of Custom Instructions

Custom Instructions can be very useful for this situation, but other more permanent measures will likely be taken in the future.
References
- https://embracethered.com/blog/posts/2023/google-bard-image-to-prompt-injection/
- https://twitter.com/AIPanic/status/1714371497492361704
Timeline
- 2023-11-01 — v1.0
- 2023-05-22 — v1.1
- 2023-05-24 — v1.2

