With Google's recent introduction of the Gemini Ultra model and the subsequent release of their Gemma models as open source, we thought it was time for some clarity on where the market stands, and to connect that to the cost discussion we often have in our work.
3-level product paradigm
In the market for large language models (LLMs), we're seeing a three-tier product strategy. The largest models, such as GPT-4 and Gemini Ultra, are built for complex reasoning. They are impressive, but they need a lot of compute and cost more to run, so they make sense in scenarios that translate into high revenues.
Mid-tier LLMs, with around 7 billion parameters, cater to business applications and some consumer uses on high-end PCs. An example is the integration of a 7B model into virtual assistant software, which makes for more natural interaction and smarter responses for tasks such as scheduling and email management directly from a user's desktop. Gemma 7B offers comparable capabilities for desktop development and gives developers more flexibility.
At the entry level, models with 2 billion parameters or less are aimed at consumer goods like phones. Running LLMs directly on PCs or phones allows for local data processing, so data doesn't have to be sent offsite for analysis. Models can be adapted on-site, tap into local resources, or connect with other applications securely, without exposing sensitive information.
Large models
These get most of the media attention. They are very capable and all have more than 70B parameters, which makes them expensive both to train and to run. Some examples are:
- GPT4 from OpenAI
- Claude2 from Anthropic
- Gemini Ultra from Google
- LLama2 from Facebook
There are more, but these are the best known. The LLama2 model from Facebook is particularly interesting because it is open source and many others build on it. A remarkable project we came across last week was Groq, which runs these large models at amazing speed. It is a hardware project that also provides access through an API. You can test it on their website and it works at 300 tokens per second. This is only possible because the inference (the technical term for generating responses or predictions) is much more efficient and needs less compute, which means the cost of hosting a model is lower as well.
Smaller models
There are hundreds of LLMs in the 7B and 2B classes. The new Gemma models were benchmarked against similar models in the same category, and especially on maths and programming tasks they score significantly better than models of the same size. Still, benchmark scores for coding and maths are well below human levels.
For other tasks, they score similarly to competitors like Mistral 7B, which is still well below the large models. That raises the question of why you would consider a model like that at all.
The answer is the same reason we were so excited about Groq, which is speed and cost. For many applications, 7B models are more than sufficient, especially when they are trained or fine-tuned for one specific task. That makes them attractive for businesses that need reliable AI without a hefty price tag. A 7B model can be a practical solution for a wide range of applications, from automated customer service to content creation, at a fraction of the cost.
The arrival of models like Gemma also shows how fast this part of the field is moving. As these smaller models improve, they get better at complex tasks and move closer to human-level performance in specific domains. That opens up areas where accuracy and efficiency matter, without the resources a larger model would require, and it makes AI more accessible across sectors in a way that scales economically.
What does it mean
Lately we have often been in discussions where operational costs are a big topic, and rightly so. Yet when we look at the developments of just the last year, we feel it is safe to say that for most applications, costs will drop significantly. That opens up AI for a much broader range of applications, without the prohibitive costs that can be a barrier today.
To benefit from these cost reductions, though, businesses need to start preparing now. That means investing in the right technology and skills, and developing a plan that fits AI into the business model in a way that lines up with where the capabilities are heading. It also means keeping up with what is happening in the field and understanding how it can be applied to real business problems.
In practice, this could mean identifying the parts of the business that stand to gain the most from AI, such as customer service, data analysis or operational efficiency. By experimenting with AI solutions now, businesses get a clearer picture of the potential impact and the operational changes needed to support it. And things can go fast. For example, Groq now offers 1M tokens for the 70B at just $0.70, compared to ±$1.20 for ChatGPT3.5, which is a 42% cost reduction. The 7B model at Groq is priced at just $0.10 per 1M tokens, which is only 8% of the ChatGPT cost.
We think these are very exciting developments that make AI accessible for so many companies already today.

