In 2012 the now-famous AlexNet beat every other model in the ImageNet competition by a wide margin, and Convolutional Neural Networks (CNNs) took over the field. Almost overnight this flavour of deep learning halved error rates compared to the best computer vision techniques of the time. It was the start of a steep climb in performance that soon approached human-level accuracy.
Now, in 2024, we are in the middle of another one. Where 2012 was about a step up in performance, this shift is about accessibility: generalist vision models that can solve a wide range of tasks, and that anyone can use. As in natural language processing, it is driven by the Transformer architecture, the same model underneath Large Language Models and ChatGPT.
The way we make computers see is about to change
The current way of doing computer vision is to collect a large number of images, label them, and train a specialised model for each task. Once trained, the model is validated on held-out data in the hope that it will cope with the real world and all its edge cases. It is cumbersome work, and it is no surprise that a whole industry of specialised computer vision companies has grown up around it, often solving tasks that are quite mundane for a human.
With the arrival of Multimodal LLMs this is about to change, and the shift is still going largely unnoticed, as Ethan Mollick pointed out in a recent tweet:

Ethan Mollick on how little attention the vision side of AI is getting.
Start experimenting with multimodal models like OpenAI’s GPT-4o, Anthropic's Claude Sonnet 3.5, Google’s PaliGemma or Tencent's YoloWorld and you can feel that something has changed. Until recently these models were only accessible to experts. Now you tell (prompt) a generalist model which vision task it needs to solve, and it does it.

Look, I trained... uhhh, prompted a pothole detector
It is still early days, so do not expect perfection
It is impressive to watch a multimodal LLM solve a vision task, but these systems still have their issues. Whether you should add one to your computer vision solution depends on the problem.
For complex vision tasks where humans struggle too and accuracy is non-negotiable, adding an LLM to your stack may not (yet) be a good idea. For more straightforward tasks where an occasional error is acceptable, experimenting with Multimodal LLMs is worth it, especially while prototyping or scaling up.
Once your solution proves useful you can still switch to training your own model, which might be cheaper and faster. Or you can distil the large model into a smaller one that runs on an edge device, for example, as shown in this blog post: LLM Knowledge Distillation: GPT-4o.
Evaluation remains a crucial factor for success
With these LLM systems you no longer have to train your own model, since companies like OpenAI have done that for you. That does not mean you can skip collecting data and doing proper evaluation.
In the current paradigm you collect a lot of images to train your model and set part of them aside for evaluation. In the new paradigm there is no training, but you still need to tune your prompts and check how well the system performs. Evaluation matters as much as it ever did. You will still need data for it, though far less than you would need to train a model from scratch.
The current generation of LLMs has annoying failure modes, and that is nothing new. It holds for Retrieval-Augmented Generation (RAG), for generating text and, in this case, for computer vision. Without good evaluations you will not get there. See for example our earlier post on Pushing a RAG Prototype to Production.
Do not underestimate the potential upside of initial success
If you have managed to solve your computer vision problem well enough with a multimodal LLM, you are in a good position. The LLM market is moving fast towards better models at lower prices (see our previous post on this, The Era of Choice in AI), and you are set up to ride that wave.
With little effort your solution will keep improving and get cheaper to run. All you need to do is keep an eye on your evaluations and adjust your prompts as newer models come out.
For years computer vision was about squeezing out more performance. From here on it is at least as much about making it simpler and more accessible for everyone.

