How to Run LLMs Locally with Ollama and Python

Install Ollama, pull an open model, and call it from Python with streaming, error handling, and a speed measurement, plus an honest look at local versus hosted.

Executive Summary: Ollama lets you run open models such as Llama 3 on your own machine and call them from Python in about ten lines. This post covers setup and usage, the benefits of privacy, offline use, and no per-token cost, and the trade-offs in memory, speed, and model quality, so you can measure before you commit.

A developer wants to summarize internal support tickets with a language model. Legal says the tickets cannot leave the company network. That one sentence rules out every hosted API, and for many teams it ends the project. It does not have to. A model that runs on a laptop never sends a byte anywhere.

Ollama is a free, open source tool that downloads open language models and serves them through a local HTTP API. The Ollama Python library is a small client for that API. Together they let you chat with a model, stream its output, and create embeddings, all on your own computer.

You need Python 3.12 and a machine with at least 8 GB of memory. No account, no API key, and no payment are involved. After the first model download, everything works offline. As with any language model, output is non-deterministic, so your replies will differ from the examples.

My position: every developer who builds with language models should have one running locally. It makes development free and private, and it teaches you what a model actually costs to run. However, local is a deployment choice with trade-offs, not a free replacement for hosted models.

How local models work

your Python script
      |  HTTP on localhost:11434
      v
Ollama server (runs in the background)
      |  loads the model file into memory, RAM or GPU
      v
model weights on disk (GGUF file, a few GB)
      |  inference by llama.cpp
      v
tokens streamed back to your script

Three terms explain most of what follows.

  • Parameters. A model’s size is its number of weights. “Llama 3 8B” has 8 billion. More parameters usually mean better answers and always mean more memory.
  • Quantization. Weights are stored with fewer bits each, commonly 4 instead of 16. The file shrinks about four times, with a modest loss of quality. This is what makes a large model fit on a laptop.
  • GGUF. The file format for quantized models used by llama.cpp, the inference engine that Ollama builds on.

You can estimate memory with one multiplication. Parameters times bits per weight, divided by 8, gives bytes. An 8 billion parameter model at 4 bits is about 4 GB, plus overhead. That matches the size of the default Llama 3 download.

Install Ollama and pull a model

Download the installer for macOS, Windows, or Linux from ollama.com and run it. Then check the command-line tool and download a model.

ollama --version
ollama pull llama3
ollama run llama3 "Explain what a virtual environment is in one sentence."

pull downloads the model once, about 4.7 GB for llama3. run loads it and prints an answer. The first reply is slow, because the model must load into memory. Later replies start quickly.

Model tag Parameters Download size Good for
phi3 3.8 billion About 2.3 GB Machines with 8 GB of memory, and simple tasks
mistral 7 billion About 4.1 GB General use on a typical laptop
llama3 8 billion About 4.7 GB General use, the default choice today
llama3:70b 70 billion About 40 GB Workstations with a large GPU or a lot of memory

The Ollama documentation recommends at least 8 GB of RAM for 7 billion parameter models, 16 GB for 13 billion, and 32 GB for 33 billion. If a model needs more memory than you have, the machine starts swapping to disk, and generation slows from words per second to words per minute.

Call Ollama from Python

python --version
python -m venv .venv
source .venv/bin/activate          # macOS and Linux
.venv\Scripts\Activate.ps1         # Windows PowerShell
python -m pip install ollama pytest

Save this complete script as local_chat.py. It handles the two failures you will meet first, and it reports generation speed.

import sys

import httpx
import ollama

MODEL = 'llama3'
HOST = 'http://localhost:11434'
SYSTEM_PROMPT = 'You are a concise assistant. Answer in two sentences at most.'


def build_client() -> ollama.Client:
    return ollama.Client(host=HOST, timeout=120.0)


def ask(client, question: str) -> tuple[str, float]:
    response = client.chat(
        model=MODEL,
        messages=[
            {'role': 'system', 'content': SYSTEM_PROMPT},
            {'role': 'user', 'content': question},
        ],
        options={'temperature': 0.2, 'num_ctx': 4096},
    )
    seconds = response['eval_duration'] / 1_000_000_000
    tokens_per_second = response['eval_count'] / seconds if seconds else 0.0
    return response['message']['content'], tokens_per_second


def main() -> int:
    question = ' '.join(sys.argv[1:]) or 'Give me one tip for naming a Python function.'
    client = build_client()
    try:
        text, speed = ask(client, question)
    except httpx.ConnectError:
        print('Cannot reach Ollama. Start the app, or run: ollama serve', file=sys.stderr)
        return 1
    except ollama.ResponseError as error:
        if error.status_code == 404:
            print(f'Model not found. Run: ollama pull {MODEL}', file=sys.stderr)
        else:
            print(f'Ollama error {error.status_code}: {error.error}', file=sys.stderr)
        return 1
    print(text)
    print(f'[{speed:.0f} tokens per second]')
    return 0


if __name__ == '__main__':
    raise SystemExit(main())
python local_chat.py

Example output. The text varies between runs, and the speed depends entirely on your hardware:

Use a verb that says what the function does, such as calculate_total, and keep it specific enough that a reader needs no comment.
[31 tokens per second]

The request looks like a call to any hosted chat API: a model name and a list of messages with roles. The difference is the address. Everything goes to localhost, so no key is needed and nothing leaves the machine. If you have used a hosted service, compare this with calling a hosted model API.

Two errors are handled explicitly. If the server is not running, the HTTP layer raises httpx.ConnectError. If you request a model you have not pulled, Ollama returns a 404, and the library raises ollama.ResponseError. The timeout is a generous 120 seconds, because the first request includes loading several gigabytes into memory.

Stream the reply

Local generation is slower than a hosted service, so streaming matters even more. Pass stream=True and print each piece as it arrives.

import ollama

MODEL = 'llama3'


def main() -> None:
    client = ollama.Client(host='http://localhost:11434', timeout=120.0)
    stream = client.chat(
        model=MODEL,
        messages=[{'role': 'user', 'content': 'Explain what a token is in three sentences.'}],
        stream=True,
    )
    for chunk in stream:
        print(chunk['message']['content'], end='', flush=True)
    print()


if __name__ == '__main__':
    main()

The setting that silently truncates your prompt

Here is the detail that causes more confusing local-model bugs than any other. Ollama’s default context window is 2,048 tokens, regardless of what the model supports. Llama 3 can handle 8,192 tokens, yet out of the box it sees only the last 2,048.

If your prompt is longer, Ollama does not raise an error. It drops the oldest part of the input. In a retrieval pipeline, that often means the system prompt and the first sources disappear, and the model answers a question it can no longer see instructions for. The reply looks plausible, and nothing in your logs says why it is wrong.

The fix is the num_ctx option, which the script sets to 4,096:

options={'num_ctx': 4096}

A larger context uses more memory, so do not simply set the maximum. My rule: decide the context size from your longest real prompt plus the reply, set num_ctx explicitly on every call, and compare response['prompt_eval_count'] with your own token estimate. If the count is lower than you expected, your prompt was cut.

Measure speed on your own hardware

Every response includes timing fields. eval_count is the number of generated tokens, and eval_duration is the time spent generating them, in nanoseconds. Their ratio is tokens per second, which the script prints.

Treat that number as the basic fact of a local deployment. On my laptop, the 8 billion parameter model produces around 30 tokens per second, which reads as fast typing. That is my own measurement, and yours will differ, so run the script. A machine without a capable GPU may produce 5 tokens per second or fewer, which is too slow for an interactive feature and acceptable for an overnight batch.

The arithmetic matters for planning. A 300-token answer at 30 tokens per second takes ten seconds. A job that summarizes 10,000 documents at that rate takes more than a day on one machine. Do this calculation before you promise anything.

Use the OpenAI-compatible endpoint

Ollama also exposes an endpoint that mimics the OpenAI chat API. Code written for the openai package can talk to a local model by changing the base URL.

from openai import OpenAI

MODEL = 'llama3'

client = OpenAI(base_url='http://localhost:11434/v1', api_key='ollama', timeout=120.0)
response = client.chat.completions.create(
    model=MODEL,
    messages=[{'role': 'user', 'content': 'Say hello in five words.'}],
)
print(response.choices[0].message.content)

The SDK insists on an API key, so pass any placeholder. Ollama ignores it. This compatibility is useful for development: point your application at the local server while you build, and at the hosted service in production. Ollama labels the feature experimental, and not every parameter is supported, so test the features you rely on.

The trade-off is prompt behavior. A prompt tuned for a large hosted model can perform poorly on an 8 billion parameter model. Switching the URL is easy. Getting equal quality is not.

Local embeddings

Ollama can also run embedding models, which turn text into vectors for search. Pull one, then call it.

ollama pull nomic-embed-text
import ollama

client = ollama.Client(host='http://localhost:11434')
result = client.embeddings(model='nomic-embed-text', prompt='Refunds take five business days.')
print(len(result['embedding']))

With a local chat model and a local embedding model, you can build semantic search with embeddings, or a complete retrieval-augmented generation pipeline, in which no document or question ever leaves your network.

Test without Ollama running

The ask function takes the client as an argument, so a test can pass a fake. The test below needs no model and no server. Save it as test_local_chat.py.

from local_chat import ask


class FakeClient:
    def __init__(self) -> None:
        self.calls: list[dict] = []

    def chat(self, **kwargs):
        self.calls.append(kwargs)
        return {
            'message': {'role': 'assistant', 'content': 'Use a clear verb.'},
            'eval_count': 50,
            'eval_duration': 2_000_000_000,
        }


def test_ask_returns_text_and_speed():
    text, speed = ask(FakeClient(), 'Any tip?')
    assert text == 'Use a clear verb.'
    assert speed == 25.0


def test_ask_sets_an_explicit_context_size():
    client = FakeClient()
    ask(client, 'Any tip?')
    assert client.calls[0]['options']['num_ctx'] == 4096
    assert [message['role'] for message in client.calls[0]['messages']] == ['system', 'user']
python -m pytest -q test_local_chat.py
..                                                                       [100%]
2 passed in 0.05s

Fifty tokens in two seconds is 25 tokens per second, which the first test checks. The second test guards the context setting, so nobody removes it by accident.

The myth: local models are free

Local models have no per-token price, and many people stop the analysis there. Nothing is billed per request, yet the costs are real.

  • Hardware. Useful speed needs a recent laptop with fast unified memory, or a GPU with enough video memory. A server GPU is a significant purchase or rental.
  • Throughput. One machine serves roughly one request at a time at full speed. Ten simultaneous users need a queue or more machines.
  • Operations. You now own model updates, monitoring, capacity, and uptime. A hosted provider does that for you.
  • Quality. A small quantized model makes more mistakes on hard reasoning than the largest hosted models. Errors have a cost too.

For a developer’s own machine, local is effectively free. For a production service with steady traffic, a hosted API is often cheaper than running your own GPUs until volume is high. Calculate it for your case.

Local versus hosted

Question Local with Ollama Hosted API
Where does the data go? Nowhere. It stays on your machine To the provider’s servers
Cost model Hardware and electricity Pay per token
Works offline Yes No
Best available quality Good for many tasks, and limited by model size The strongest models available
Speed Depends on your hardware Usually fast and consistent
Concurrency Limited by one machine Scales with your rate limit
Rate limits None Yes
Who runs it? You The provider

The hosted side keeps moving quickly. OpenAI announced GPT-4o this week, and open models keep improving too. Keep model names in constants, and re-evaluate your choice a few times a year.

How real systems use local models

  • Development and CI. Teams run a small local model while they build and test, so that everyday work costs nothing and needs no shared key.
  • Private data processing. Summarizing, classifying, and redacting sensitive documents happens on machines inside the company network.
  • Hybrid routing. Simple, high-volume tasks go to a local model, and hard or rare ones go to a hosted model.
  • Offline and edge use. Field devices and air-gapped environments run a small model where no network exists.
  • One client interface. Application code depends on a small wrapper, so that the model behind it can be local or hosted by configuration.

We once hit a bug when a document summarizer that worked well on short test files produced confident summaries of the wrong content on real reports. The prompts were around 6,000 tokens, and the local server was running with its default context of 2,048. It kept only the tail of each document and discarded our instructions. No error appeared anywhere. Setting num_ctx explicitly fixed it, and we added a check that compares prompt_eval_count with our own token estimate.

Choosing local or hosted: a decision framework

  1. Is the data allowed to leave your network? If not, run locally. The decision is made.
  2. Is the task simple and repetitive? Classification, extraction, and short summaries suit an 8 billion parameter model.
  3. Does it need strong reasoning or long context? Use a large hosted model, or accept lower quality.
  4. How many tokens per second do you get? Measure, then multiply by your workload. If the result is too slow, you need better hardware or a hosted API.
  5. How many users at once? One machine handles little concurrency. Interactive products with many users usually belong on a hosted service.
  6. Who will operate it? If nobody owns the server, choose hosted.

When NOT to run models locally

  • You need the best available quality. For hard reasoning, long documents, and subtle writing, the largest hosted models are still ahead of anything that fits on a laptop.
  • You serve many users at the same time. A single machine queues requests. Matching a hosted service’s concurrency means running a GPU cluster.
  • Your hardware is too small. A model that barely fits in memory runs at a crawl. With less than 8 GB of RAM, a hosted API will serve you better.

Common mistakes

  • Leaving the default context size. Prompts longer than 2,048 tokens are cut without warning. Instructions and early content vanish.
  • Using a short timeout. The first request loads the model, which can take many seconds. A five-second timeout fails every cold start.
  • Pulling a model that does not fit. A 70 billion parameter model on a 16 GB laptop swaps to disk and produces a word every few seconds.
  • Exposing the port to the network. Ollama has no built-in authentication. Binding it to all interfaces lets anyone on the network use your machine. Keep it on localhost, or put an authenticated proxy in front.
  • Reusing hosted-model prompts unchanged. Small models follow long, subtle instructions less reliably. Shorten prompts, and test again.
  • Comparing only price per token. Hardware, operations, and error rates are costs too. A free model that is wrong more often is not free.

Key takeaways

  • Ollama downloads open models and serves them at http://localhost:11434.
  • The ollama Python library calls it with client.chat(), and stream=True streams the reply.
  • Estimate memory as parameters times bits per weight, divided by 8.
  • Set num_ctx on every call, because the default of 2,048 tokens truncates silently.
  • Compute tokens per second from eval_count and eval_duration, and plan from that number.
  • The OpenAI-compatible endpoint lets existing code target a local model.
  • Local means private and offline, at the cost of hardware, speed, and peak quality.

FAQ

How do I run an LLM locally with Python?

Install Ollama, run ollama pull llama3, install the ollama Python package, and call ollama.Client().chat() with a model name and a list of messages.

How much RAM do I need to run Llama 3 locally?

The 8 billion parameter version needs at least 8 GB of RAM, and 16 GB is more comfortable. The 70 billion parameter version needs roughly 40 GB or more of memory.

Is Ollama free?

Yes. Ollama is open source and free to use, and the models it runs have no per-token charge. You pay for the hardware and electricity.

Can I use the OpenAI Python library with Ollama?

Yes. Create the client with base_url='http://localhost:11434/v1' and any placeholder API key. Ollama provides an OpenAI-compatible chat endpoint, which it labels experimental.

Are local LLMs as good as hosted models?

For simple tasks such as classification, extraction, and short summaries, small local models often do well. For hard reasoning and long documents, the largest hosted models are still stronger.

Measure your hardware, then decide

Running a model locally turns an abstract service into something you can time and inspect. Ten lines of Python get you a private, offline assistant. Whether it belongs in production depends on numbers you can now collect yourself: tokens per second, memory, and how often the answers are right.

Rule of thumb: keep data local when you must, keep tasks simple when you can, and never trust a context window you did not set yourself.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *