The webfetch tool fetches remote URLs and returns their content with intelligent
extraction designed for documentation, static pages, and structured text. It
provides caching, llms.txt probing, binary content handling, and optional
secondary-model summarization.
webfetch is already a built-in tool in OpenCode. This plugin replaces it with
an enhanced version — the implementation lives in src/tools/smartfetch/, and
the tool is registered under the same webfetch name to override the default.
| Parameter | Type | Default | Description |
|---|---|---|---|
url |
URL (string) | required | The URL to fetch. Must be a valid HTTP/HTTPS URL. |
format |
"text" | "markdown" | "html" |
"markdown" |
Output format for the fetched content. |
timeout |
number | 30 |
Timeout in seconds (max 120). |
prompt |
string | optional | An extraction task for the secondary model to run against the fetched content (see Secondary Model). |
extract_main |
boolean | true |
Extract main content from HTML using Mozilla Readability. When disabled, returns the full page body. |
prefer_llms_txt |
"auto" | "always" | "never" |
"auto" |
Prefer /llms.txt or /llms-full.txt over the page itself. "auto" probes only for docs-like domains (readthedocs, gitbook, netlify, vercel, etc.). |
include_metadata |
boolean | true |
Include YAML frontmatter with fetch metadata (status code, content type, charset, redirect chain, cache info, etc.). |
save_binary |
boolean | false |
Save binary payloads (images, PDFs, audio, video) to disk under the system temp dir. When disabled, binary content reports metadata-only. |
Returns the fetched content in the requested format. When include_metadata
is enabled (default), the response is prefixed with YAML frontmatter containing
metadata about the fetch:
---
requested_url: "https://example.com/docs"
final_url: "https://example.com/docs"
canonical_url: "https://example.com/docs"
status_code: 200
source_content_type: "text/html"
source_kind: "html"
title: "Documentation"
headings:
- "Getting Started"
- "API Reference"
used_llms_txt: false
extracted_main: true
redirect_chain: []
upgraded_to_https: true
cache_hit: false
word_count: 1420
quality_signals: []
truncated: false
---
The quality_signals field flags potential issues:
very_short_content — fewer than 60 wordspossible_paywall — content matches paywall/login keywordshigh_boilerplate_ratio — large HTML-to-text ratio without Readability extractionBinary responses (images, PDFs, audio, video) return metadata about the file:
Content-Disposition or URL path)image, audio, video, pdf, binary)Two modes:
save_binary, 10 MiB with it). Reports size and type without the body.save_binary=true, the binary is written to
<tmpdir>/opencode-smartfetch/<filename> and the response includes the
filesystem path.When a cross-origin redirect is blocked by policy, the response explains which URL was attempted and provides the redirect URL so you can fetch it directly.
When a prompt parameter is supplied, webfetch can route the fetched content
through a secondary (cheaper) model for focused extraction. This lets you ask
questions like "summarize this page" or "extract the code examples" in one step.
How it works:
Which model is used (in priority order):
webfetch.model (dedicated — highest priority, supports array for fallback)small_model from the OpenCode configuration (opencode.json / opencode.jsonc)explorer agent modellibrarian agent modelThe secondary model is called only when all of these are true:
prompt parameter is providedIf the secondary model fails (timeout, error, empty response), webfetch
returns the raw fetched content as a graceful fallback.
Fetches are cached in memory with an LRU cache (50 MiB max, 15-minute TTL).
The cache key includes the URL plus behavior-affecting options (extract_main,
prefer_llms_txt, save_binary), so changing these re-fetches the URL.
Revalidation: Cache entries with ETag or Last-Modified headers support
conditional revalidation. When a stale entry exists, webfetch sends
If-None-Match / If-Modified-Since headers. A 304 Not Modified response
refreshes the TTL without re-downloading.
llms.txt validation: Cached llms.txt results are validated — if the
cached entry doesn't actually look like an llms.txt response (wrong path,
HTML content, login page), it is evicted and re-fetched.
For documentation sites, webfetch probes for /llms-full.txt then /llms.txt
before falling back to the page itself.
Probing behavior depends on the prefer_llms_txt parameter:
"auto" (default) — probes only when the domain looks documentation-adjacent
(suffixes like .readthedocs.io, .gitbook.io, docs.rs; prefixes like
docs., developer., dev., wiki.)"always" — always probes; fails with a message if neither llms.txt variant
exists"never" — skips probing entirelyThe probe respects cross-origin redirect policy (same origin only). If the
llms.txt response is HTML or a login page, the probe is rejected.
webfetch follows up to 10 redirects per request, but only within same-origin
scopes. Cross-origin redirects are blocked and the caller is instructed to
fetch the new URL directly.
For URLs entered as http://, webfetch first tries https:// and falls
back to http:// if the HTTPS attempt fails (connection error, blocked
redirect, or non-2xx status).
Content type detection follows this flow:
image/*, audio/*, video/*,
application/pdf, application/zip, application/octet-stream) are
treated as binary.application/octet-stream and known text types are re-examined — the
first 2 KiB is scanned for null bytes and non-printable characters to
distinguish text from binary.text/html for better content extraction.Set webfetch.enabled to false to skip registering the enhanced version and
use OpenCode's built-in webfetch instead:
{
"webfetch": {
"enabled": false
}
}
The webfetch.model option sets a dedicated model (or array of fallback
models) for secondary-model summarization. Takes priority over all other model
resolution sources. Accepts the same format as agent model configs:
{
"webfetch": {
"model": "openai/gpt-4o-mini"
}
}
Multiple fallback models in priority order:
{
"webfetch": {
"model": ["openai/gpt-4o-mini", "anthropic/claude-3-haiku"]
}
}
With optional variant:
{
"webfetch": {
"model": [
"openai/gpt-4o-mini",
{ "id": "anthropic/claude-3-haiku", "variant": "low-latency" }
]
}
}
Each entry is tried in turn; the first to return usable text is used.
The secondary model is resolved from these sources (in priority order):
webfetch.model (dedicated — highest priority, supports array for fallback)small_model in the OpenCode config (opencode.json / opencode.jsonc at
project or user level)agents.explorer.model configagents.librarian.model configExample opencode.jsonc:
{
"small_model": "openai/gpt-4o-mini"
}
Or in the plugin's opencode.json preset or project config:
{
"agents": {
"explorer": { "model": "anthropic/claude-3-haiku" },
"librarian": { "model": "openai/gpt-4o-mini" }
}
}
The webfetch permission can be configured in the plugin's permission rules.
See Configuration for details.
The tool is registered under the name webfetch in src/index.ts, which
overrides OpenCode's built-in webfetch when this plugin is active.
The enhanced webfetch lives in src/tools/smartfetch/ (the internal module is
named "smartfetch", while the public tool name is webfetch). It is composed of
these modules:
| Module | Responsibility |
|---|---|
tool.ts |
Entry point — permission prompts, cache lookup, llms.txt preference logic, binary-vs-text branching, metadata emission, secondary-model integration |
network.ts |
URL normalization, redirect policy, charset/body decoding, header extraction, llms.txt probing, HTTP fetch with HTTPS upgrade fallback |
utils.ts |
HTML extraction (Mozilla Readability + Turndown), heading cleanup, markdown/text cleaning, frontmatter generation, quality signal detection |
cache.ts |
LRU cache keyed by URL + behavioral options, conditional revalidation, canonical URL aliasing, llms result invalidation |
binary.ts |
Binary content persistence to disk, MIME-to-extension mapping, safe filename allocation |
secondary-model.ts |
Dedicated webfetch/small_model config resolution, temporary session creation, content truncation, model fallback chain |
constants.ts |
Timeouts, size limits, docs domain heuristics, binary MIME prefixes, tool description |