In a previous article I covered getting GPTel running with LM Studio and a local Qwen model — a great setup when privacy and offline access matter. But the Emacs AI ecosystem has grown considerably, and there is now a whole family of packages targeting different parts of a coding workflow. This article picks up where that one left off: connecting GPTel to Claude's API for heavy reasoning work, introducing Ellama as a batteries-included alternative for local tasks, adding inline completions with copilot.el or codeium.el, and tying it all together into a layered workflow where each tool handles what it does best.
Updated October 2026: current Claude model names, a fix for the thinking configuration (the old budget_tokens form is now rejected by the API), and a new section on the emacs-devel debate about whether Emacs packages should point users at cloud models at all.
The Layered Approach
The key insight is that no single AI tool covers every use case well. Instead of picking one and compromising everywhere, think in layers:
- Inline completions — ghost-text suggestions as you type (copilot.el, codeium.el)
- Chat and reasoning — ask questions, explain code, generate functions (GPTel + Claude)
- Routine local tasks — summarise, review, translate, improve grammar (Ellama + Ollama)
- Multi-file agentic edits — large refactors across an entire codebase (aidermacs)
Each layer uses a different model tier. Inline completions need a fast, cheap model. Chat and reasoning work benefits from a capable cloud model. Local tasks can stay on-device. Agentic edits often use two models — one for reasoning, one for writing code. You only pay for cloud inference when it genuinely earns its cost.
GPTel with Claude (Anthropic API)
The GPTel package already supports Anthropic natively since version 0.9. If you have it installed, adding Claude as a backend takes a few lines of Elisp.
Storing the API Key Securely
Never hardcode an API key in your config. Store it in ~/.authinfo (or ~/.authinfo.gpg for GPG encryption):
machine api.anthropic.com login apikey password sk-ant-YOUR-KEY-HERE
GPTel reads this automatically when gptel-api-key is left at its default value. No key in your init file, no accidental commits to a public dotfiles repo.
Basic Configuration
(use-package gptel
:ensure t
:config
;; Register Claude as the default backend and model. :models is listed
;; explicitly because gptel's built-in model list can lag a new release.
(setq gptel-backend (gptel-make-anthropic "Claude"
:stream t
:models '(claude-sonnet-5-5 claude-opus-5-5))
gptel-model 'claude-sonnet-5-5))
The explicit :models list is there for a practical reason: the gptel build from early October 2026 knows claude-opus-5-5 but not yet claude-sonnet-5-5, so without it the newest Sonnet doesn't show up in the model menu.
Thinking and Effort
Claude can reason through a problem before answering, which helps most on algorithmic questions, debugging, and architecture decisions. The first version of this article enabled that with :thinking (:type "enabled" :budget_tokens 4096). Current Claude models reject that form with a 400 error: the fixed thinking budget is gone. Thinking is now adaptive, meaning the model decides how much to reason, and you steer how hard it works with an effort level. Register a separate backend for hard problems:
(gptel-make-anthropic "Claude-deep"
:stream t
:models '(claude-opus-5-5)
:request-params '(:thinking (:type "adaptive")
:output_config (:effort "high")
:max_tokens 16000))
gptel merges :request-params into the request body, so these go to the API as-is. Effort runs from low through medium, high and xhigh to max. Switch to this backend for hard problems with M-x gptel-menu (C-c RET) → Backend → Claude-deep, and back to the standard one for quick questions where the extra latency isn't worth it.
Core Keybindings
(use-package gptel
:ensure t
:bind (("C-c l" . gptel) ; open/switch to chat buffer
("C-c C-<" . gptel-send) ; send region or buffer
("C-c RET" . gptel-menu))) ; full options menu
The most useful interaction pattern: select a region of code, call gptel-send (C-c C-<), and the response is inserted directly into the buffer below your selection. No copy-paste, no context switching.
Using GPTel Inside Org Mode
GPTel has first-class Org mode support that makes it significantly more useful for long-running work:
;; Limit context to the current Org heading — stops the whole document
;; being sent as context for every message
M-x gptel-org-set-topic
;; Branch conversations per heading — each subtree gets its own thread
M-x gptel-org-branching-context
;; Save backend/model as Org properties so different headings
;; use different models automatically
M-x gptel-org-set-properties
A practical pattern: keep a dev-notes.org file with one heading per feature or bug. Each heading becomes its own Claude conversation with full branching context. Your conversation history lives in a plain text file, version-controlled alongside your code.
Ellama: Batteries Included for Local Tasks
Ellama takes a different philosophy from GPTel. Where GPTel is a minimal, composable chat layer you build workflows on top of, Ellama ships with dozens of ready-made commands for specific tasks. It is built on the llm package rather than GPTel's own backend code, and it defaults to Ollama for local inference.
Installation and Setup
;; Install Ollama first: https://ollama.com
;; Then pull a model: ollama pull llama3.2
(use-package ellama
:ensure t
:init
(setopt ellama-language "English"
ellama-keymap-prefix "C-c e")
:config
(require 'llm-ollama)
(setopt ellama-provider
(make-llm-ollama
:chat-model "llama3.2"
:embedding-model "nomic-embed-text")))
With C-c e as the prefix, you get a transient menu of commands. The most useful for coding:
| Command | Keybinding | What it does |
|---|---|---|
ellama-code-review | C-c e c r | Reviews selected code and suggests improvements |
ellama-code-complete | C-c e c c | Completes selected code in context |
ellama-code-add | C-c e c a | Adds code based on a description you provide |
ellama-code-edit | C-c e c e | Edits code based on an instruction |
ellama-summarize | C-c e s | Summarises selected text or buffer |
ellama-improve-grammar | C-c e i g | Fixes grammar in selected text |
Multiple Providers
One of Ellama's strengths is routing different tasks to different models. Configure separate providers and switch between them:
(require 'llm-ollama)
(require 'llm-openai)
(setopt ellama-providers
'(("local-llama" . (make-llm-ollama :chat-model "llama3.2"))
("local-coder" . (make-llm-ollama :chat-model "qwen2.5-coder:7b"))
("claude" . (make-llm-openai-compatible
:url "https://api.anthropic.com/v1/"
:chat-model "claude-sonnet-5-5"
:key (auth-source-pick-first-password
:host "api.anthropic.com")))))
Switch providers with M-x ellama-provider-select. Use the local coding model for routine completions, switch to Claude when you need stronger reasoning.
Session Persistence
Ellama auto-saves conversation sessions and lets you load, rename, or delete them:
M-x ellama-load-session ; resume a previous conversation
M-x ellama-rename-session ; give the current session a meaningful name
Sessions are stored as Org files in ellama-sessions-directory, so they are plain text and searchable.
Inline Completions: copilot.el vs codeium.el
Both GPTel and Ellama are on-demand tools — you invoke them explicitly. Inline completion packages work differently: they watch what you type and offer ghost-text suggestions in real time, similar to GitHub Copilot in VS Code.
copilot.el
copilot.el connects to GitHub's official Copilot language server. As of early 2025 it uses the @github/copilot-language-server Node package directly (dropping the previous copilot.vim dependency). It requires Node.js 22+ and a GitHub Copilot subscription (a free tier is available with limited completions).
(use-package copilot
:ensure t
:hook (prog-mode . copilot-mode)
:bind (:map copilot-completion-map
("<tab>" . copilot-accept-completion)
("TAB" . copilot-accept-completion)
("C-TAB" . copilot-accept-completion-by-word)
("C-n" . copilot-next-completion)
("C-p" . copilot-previous-completion)))
After installing, run M-x copilot-install-server once, then M-x copilot-login to authenticate with GitHub. Ghost text appears as you type; TAB accepts the full suggestion, C-TAB accepts word by word.
The standout feature added in 2025 is Next Edit Suggestions (NES): Copilot predicts where you will need to make related edits elsewhere in the file based on your recent changes, and surfaces those suggestions proactively. It is genuinely useful during refactoring.
codeium.el
codeium.el is the free alternative. It uses Codeium's AI backend and integrates via completion-at-point-functions, plugging naturally into company-mode or corfu rather than using a ghost-text overlay.
(use-package codeium
:ensure t
:init
(add-to-list 'completion-at-point-functions #'codeium-completion-at-point)
:config
;; Run M-x codeium-install once to download the binary
;; Run M-x codeium-auth to authenticate
(setq codeium-mode-line-enable
(lambda (api) (not (memq api '(CancellableGetCompletions Heartbeat AcceptCompletion))))))
Which to choose: If you already pay for GitHub Copilot, copilot.el is the obvious pick — NES alone justifies it. If you want free inline completions with no subscription, codeium.el works well and integrates more naturally with Emacs's completion framework.
aidermacs: Agentic Multi-File Editing
GPTel and Ellama work well at the function or file level. For larger tasks — migrating an API, refactoring across a module, or making a change that touches a dozen files — aidermacs is the right tool. It wraps Aider, the terminal-based AI pair programmer, with an Emacs-native interface.
Setup
# Install Aider
pip install aider-chat
# Set your API key (or use ~/.authinfo as above)
export ANTHROPIC_API_KEY=sk-ant-YOUR-KEY-HERE
;; aidermacs is on NonGNU ELPA
(use-package aidermacs
:ensure t
:bind (("C-c a" . aidermacs-transient-menu))
:config
;; Use Claude Sonnet for reasoning, Haiku for code generation (faster + cheaper)
(setq aidermacs-extra-args
'("--model" "anthropic/claude-sonnet-5-5"
"--editor-model" "anthropic/claude-haiku-4-5")))
The --model / --editor-model split is Aider's Architect Mode: a capable model handles reasoning and planning, a faster/cheaper model does the actual code writing. This cuts costs significantly on large tasks without sacrificing quality.
Workflow
- Open any file in your project, call
C-c ato open the aidermacs transient menu. - Add files to the Aider context with
aidermacs-add-current-fileoraidermacs-add-files-in-dir. - Describe what you want in plain English: "Extract the authentication logic from UserController into a separate AuthService class and update all call sites."
- Aider makes the changes. Magit's
auto-revert-moderefreshes your buffers automatically. - Review the diff in Magit or with aidermacs's built-in ediff integration. Accept, reject, or ask for revisions.
aidermacs also supports Architect Mode for TDD: describe the desired behaviour, let the reasoning model write tests first, then let the editor model write the implementation to pass them.
GPTel Tool Use
GPTel 0.9+ supports tool use (function calling), which lets Claude interact with your Emacs environment directly — reading buffers, running shell commands, creating files. This is the foundation for building lightweight agents without leaving Emacs.
;; A tool that reads an Emacs buffer and returns its contents
(gptel-make-tool
:name "read_buffer"
:description "Return the contents of an Emacs buffer"
:function (lambda (buffer)
(with-current-buffer buffer
(buffer-substring-no-properties (point-min) (point-max))))
:args (list '(:name "buffer"
:type string
:description "The name of the buffer to read")))
;; A tool that runs a shell command and returns stdout
(gptel-make-tool
:name "run_command"
:description "Run a shell command and return its output"
:function (lambda (command)
(shell-command-to-string command))
:args (list '(:name "command"
:type string
:description "The shell command to run")))
With tools like these registered, you can ask Claude: "Look at the test output in the *compilation* buffer and suggest what is wrong with the failing test in src/auth.php" — and it will read both buffers itself rather than waiting for you to paste content.
GPTel also supports MCP (Model Context Protocol) via mcp.el, which lets you connect any MCP-compatible tool server:
(require 'gptel-integrations)
;; Then: M-x gptel-mcp-connect to register an MCP server
Putting It Together: A Practical Configuration
Here is a complete configuration that wires up all four layers:
;;; AI layer 1 — Inline completions (free, always on)
(use-package codeium
:ensure t
:init
(add-to-list 'completion-at-point-functions #'codeium-completion-at-point))
;;; AI layer 2 — Chat and reasoning (GPTel + Claude)
(use-package gptel
:ensure t
:bind (("C-c l" . gptel)
("C-c C-<" . gptel-send)
("C-c RET" . gptel-menu))
:config
;; Deeper-reasoning backend for hard problems
(gptel-make-anthropic "Claude-deep"
:stream t
:models '(claude-opus-5-5)
:request-params '(:thinking (:type "adaptive")
:output_config (:effort "high")
:max_tokens 16000))
;; Default to the standard Claude backend
(setq gptel-backend (gptel-make-anthropic "Claude"
:stream t
:models '(claude-sonnet-5-5 claude-opus-5-5))
gptel-model 'claude-sonnet-5-5))
;;; AI layer 3 — Routine local tasks (Ellama + Ollama)
(use-package ellama
:ensure t
:init
(setopt ellama-keymap-prefix "C-c e")
:config
(require 'llm-ollama)
(setopt ellama-provider
(make-llm-ollama
:chat-model "llama3.2"
:embedding-model "nomic-embed-text")))
;;; AI layer 4 — Agentic multi-file edits (aidermacs)
(use-package aidermacs
:ensure t
:bind ("C-c a" . aidermacs-transient-menu)
:config
(setq aidermacs-extra-args
'("--model" "anthropic/claude-sonnet-5-5"
"--editor-model" "anthropic/claude-haiku-4-5")))
With this setup, the decision of which tool to reach for becomes intuitive:
- Typing code → codeium inline suggestions appear automatically
- Need to understand or fix something →
C-c C-<sends to Claude via GPTel - Quick local task (summarise, review, grammar) →
C-c etransient menu via Ellama - Large refactor across multiple files →
C-c aopens aidermacs
Tips for Effective AI Coding in Emacs
Give context explicitly
By default GPTel sends only what is in the current buffer up to point. Use gptel-add to attach additional files as context before sending:
M-x gptel-add ; prompts for a file or buffer to add to the context
Use Org mode as a scratchpad
Open a dedicated Org file for AI conversations. Use headings to separate topics, gptel-org-branching-context to keep each thread independent, and Org's folding to manage long conversations. The file is plain text and stays in your project's git history.
Keep prompts in your config
If you find yourself typing the same instruction repeatedly, turn it into a named gptel directive:
(setq gptel-directives
'((default . "You are a helpful assistant.")
(code-review . "Review the following code for bugs, performance issues, and style. Be concise.")
(php . "You are an expert PHP developer. Suggest idiomatic PHP 8.3+ solutions.")))
Switch directive with C-c RET → System.
Always review AI diffs in Magit
Whether you use GPTel, Ellama, or aidermacs, establish the habit of reviewing changes in Magit before staging. AI-generated code can look correct but contain subtle errors — diffing keeps you in control.
Pick the right model for the task
Spending tokens on Claude Opus for a grammar correction wastes money. Use a local model via Ellama for routine tasks, reserve the capable cloud model for reasoning-heavy work. The layered setup above makes this the natural default rather than a conscious decision each time.
Should an Emacs Package Point You at a Cloud Model at All?
Most of this article assumes you're comfortable sending code to a cloud API. In late September 2026 that assumption turned into an argument on the emacs-devel mailing list.
On 23 September, in a thread titled "LLM packages in ELPA and SaaSS: a consistency question", Jean Louis listed packages that he said steer users toward hosted LLM services, naming "minuet, codex-ide, gptel, and aidermacs", and wrote that "GNU policy states that a package must not lead users to SaaSS". SaaSS is the FSF's term for "Service as a Software Substitute". Four days later he posted a "Full GNU Standards Compliance Report for llm Package", going through seven modules of the llm library that Ellama is built on, from llm-claude.el and llm-openai.el to llm-gemini.el and llm-vertex.el. He based it on the References section of the GNU Coding Standards, which says a GNU program should not "recommend, promote, or grant legitimacy to the use of any nonfree program", and applies that to anything online that "invites them to surrender control of their own computing to that site's server".
Richard Stallman's reply was narrower than a blanket rule. He wrote that "there are certainly situations where package P provides access to certain specific services and that goes against the moral criteria in the GNU Coding Standards", and that when a package "recommends or directly uses a particular program or service", the reference "might be unacceptable -- based on the facts of the situation." That's where the thread stood when I read it at the end of September.
Whatever you make of the policy question, it's a fair one to put to your own config. The gptel setup I run day to day talks only to a local Ollama server:
(setq gptel-backend
(gptel-make-ollama "Ollama"
:host "127.0.0.1:11434"
:models '(qwen3-coder:30b)
:protocol "http"
:key "not-needed"))
(setq gptel-use-curl t
gptel-stream nil)
The LM Studio backend from the previous article is still in that file, commented out. A separate, unpushed commit in another checkout of the same config adds a Mistral API backend next to Ollama, with the key read from an environment variable, but that isn't the configuration I run. The layered setup in this article works either way: the layers are independent, so a local-only setup is the same layout with the cloud entries left out.
Summary
The Emacs AI toolkit in 2026 is mature and genuinely useful for professional development — but it works best when you treat it as a set of composable layers rather than a single tool:
- codeium.el or copilot.el for always-on inline completions as you type
- GPTel + Claude for on-demand chat, code generation, and extended reasoning — with Org mode integration for persistent conversation history
- Ellama + Ollama for fast, private, offline routine tasks with a menu of ready-made commands
- aidermacs for multi-file agentic edits and large refactors, with ediff review and Architect Mode
Each tool handles what it does best. You stay in Emacs throughout. The result is an AI-augmented workflow that actually fits how Emacs developers work — keyboard-driven, buffer-centric, and composable.