Multimodal AI
Can AI understand my screenshot too, not just the text I send it?
Can AI understand my screenshot too, not just the text I send it?
This example takes text and images as input and returns text. Supporting image input does not mean it can output images.
A mode is a form of information:Text, images, audio, and video are common modes. Support for one does not mean support for all.
The same mode can be on either side:Audio can be input for transcription or output that reads an answer aloud.
Check the specific product:Supported inputs, outputs, and file types vary by model and product.
This screenshot shows the checkout page on a phone. The payment button should be at the bottom, but it is not visible. First, use only the screenshot to identify any visible overlap or layout clues, and say when the image is not enough to confirm something. Then inspect the mobile layout in the project, fix the cause, and verify at the same screen size that the button is visible and clickable.