Skip to main contentSkip to navigation
AI basics · F

Foundation model

A foundation model (base model) is a large AI model pre-trained on vast, broadly spread amounts of data that can then be adapted for many different downstream tasks. Well-known examples are large language models such as GPT, Claude and Gemini. Foundation models form the technological basis of numerous modern AI applications, from chatbots through text generation to AI agents.

Also known as: base model, foundation model, base AI model

What is a foundation model?

A foundation model, often called a base model, is a particularly large AI model pre-trained in a laborious process on enormous, thematically broad amounts of data. Rather than being developed for a single, narrowly defined task, the model learns general patterns, relationships and representations that carry over to a great many use cases.

The term was coined to describe a new generation of models serving as a common basis for many different applications. Instead of training a separate model from scratch for every task, companies use a base model pre-trained once and adapt it deliberately. This decoupling of general pre-training from specific adaptation is foundation models' central feature.

Characteristic of foundation models is their sheer size: they consist of an enormous number of parameters and are trained on amounts of data a single company could hardly assemble or process itself. Out of that extensive training comes a remarkable versatility. Models originally trained only to continue text can afterwards translate, summarise, program or argue without ever being programmed explicitly for those individual tasks. That very transferability to previously unseen tasks is what makes foundation models so valuable and distinguishes them from classic, narrowly specialised AI systems of earlier generations.

How does a foundation model work?

Foundation models are usually based on the Transformer architecture and are, as part of what is called self-Supervised learning trained. The model learns from large amounts of unstructured data, text from the internet for instance, by predicting the next word in a sentence. In this way it acquires a broad understanding of language, facts and relationships without all the training data having to be labelled by hand.

After pre-training, you have a versatile model. It is adapted further for specific use cases, for instance through Fine-tuning on specific data, through Prompt engineering or by connecting external knowledge sources via RAG (retrieval-augmented generation). That turns a general base model into a specialised solution.

Many foundation models also go through a subsequent alignment phase in which they are trained with human feedback to give helpful, safe answers that match users' expectations. This step, often called Reinforcement learning from human feedback, makes sure that a raw, purely statistical Language model becomes an assistant fit for practice. Only the interplay of broad pre-training, deliberate alignment and subsequent adaptation to the use case makes modern foundation models so capable in everyday use.

Which foundation models are there?

The best-known foundation models are large language models (LLMs). These include the GPTseries from OpenAIthat Claude models of Anthropic and the Geminimodels from Google. These models process and produce natural language and underpin many well-known AI assistants and generative applications.

Foundation models are not limited to text, though. There are also multimodal base models that additionally process images, audio or video, and specialised models for program code or scientific applications. Alongside large providers' proprietary models there are open model families such as Llama or Mistralthat give organisations more control and room to adapt.

What benefits do foundation models offer companies?

The biggest advantage lies in reusability. One single pre-trained base model can serve as the foundation for many different applications, from customer communication through text generation to complex AI agents. Companies do not have to carry out the laborious, costly pre-training themselves but can build on existing models.

The barriers to entry for AI projects thus fall considerably. Through APIs, capable models can be integrated into existing systems in a short time. At the same time, adaptability through fine-tuning or company-specific knowledge lets solutions be tailored precisely to the use case at hand.

An advantage in speed comes on top: because the laborious foundational training is already done, prototypes and pilot projects can be realised considerably faster than before. Companies can test ideas quickly, validate them and scale them on success, without lengthy development cycles. They also benefit continuously from the model providers' progress, since new, more capable model generations often become available through the same interfaces. That makes foundation models a future-proof basis for a broad AI strategy.

What challenges and limits are there?

Foundation models are capable but not infallible. They can produce factually wrong content, so-called hallucinations, and may reflect distortions from their training data. Their knowledge is also limited to the training cut-off, so current information often has to be added through further methods such as RAG.

There are also requirements on data protection, governance and compliance, for instance under the EU AI Act. Anyone wanting to use foundation models responsibly therefore needs not only the right technology but well-considered processes for quality assurance and risk assessment. At Elisabit we help companies choose the right base model, integrate it safely and tailor it to the specific use case, so real value emerges from a general model.

Frequently asked questions

What is the difference between a foundation model and an LLM?

A large language model (LLM) is a particular kind of foundation model, specialised in language. The term foundation model is broader and also covers multimodal modelsthat handle images or audio, for instance. Every LLM is a foundation model, but not every foundation model is an LLM.

Do companies have to train their own foundation model?

In most cases no. Pre-training a foundation model is extremely demanding and expensive. Companies instead use existing base models through APIs or open models and adapt them to their needs with fine-tuning, prompt engineering or RAG.

Are foundation models always language models?

No. Language models are the best-known examples, but there are also multimodal foundation models for images, audio and video and specialised models for program code among others. What is decisive is the principle of broad pre-training with subsequent adaptability.

How do you adapt a foundation model to a use case?

There are several routes: prompt engineering steers the model through well-crafted instructions, RAG brings in current or company-owned knowledge, and fine-tuning trains the model further on specific data. These methods are often combined.

Put AI to work for your business?

We help you integrate artificial intelligence into your processes, your marketing and your website — strategically and securely.

Request a project

Stefan

Your contact

Stefan

I look forward to hearing about your project and finding the best solution together.