README file from
GithubHybrid Chat
Hybrid Chat is a desktop-only Obsidian 1.13+ plugin that provides federated, provider-neutral RAG chat over existing Obsidian Hybrid Search (OHS) services.
It does not create a vector index, read OHS SQLite databases, or duplicate indexing. Retrieval is read-only: the plugin calls only the OHS MCP search and read tools. The user-invoked copy actions write Markdown to the system clipboard; the plugin never reads clipboard contents and does not create or modify vault notes.
Dependencies
Hybrid Chat is a thin client over existing services and depends on:
- Obsidian 1.13.0 or newer, desktop only (macOS, Windows, or Linux). The plugin does not run on mobile or in the Obsidian web app.
- One Obsidian Hybrid Search (OHS) endpoint per vault you want to search — a running OHS Streamable HTTP MCP endpoint (default
http://127.0.0.1:3939/mcp). Hybrid Chat does not bundle or start OHS; it only queries the endpoint'ssearchandreadtools read-only. - One OpenAI-compatible chat provider — any server exposing
POST /v1/chat/completions, for example a local LM Studio server (default base URLhttp://127.0.0.1:1234/v1) or a remote HTTPS endpoint. Remote providers must use HTTPS; plain HTTP is accepted only on loopback hosts. Provider API keys are stored as ObsidianSecretStoragesecrets, never in plugin data. - Building from source only: Node.js 24 and npm.
Installation
- Prepare the prerequisites (once per machine):
- In each vault you want to chat with, install and configure the OHS plugin and start its Streamable HTTP MCP endpoint — see the OHS repository for installation and endpoint setup.
- Run an OpenAI-compatible model server (for example, LM Studio's local server) and note its base URL, model ID, and API key if it has one.
- Install Hybrid Chat: in Obsidian, open Settings → Community plugins → Browse and install Hybrid Chat, or manually download
main.js,manifest.json, andstyles.cssfrom the latest release into<vault>/.obsidian/plugins/hybrid-chat/. - Enable the plugin and open the chat via the ribbon icon or the Hybrid Chat: Open chat command.
- Configure it: under Settings → Hybrid Chat, register one OHS endpoint per vault (stable ID, display name, exact Obsidian vault name, endpoint URL, timeout) and create an OpenAI-compatible provider profile (base URL, model, API-key secret); then select the active profile. See Configuration for the full option list.
Architecture
- The sidebar selects the current vault, explicitly selected vaults, or every enabled vault, and restores the last-used scope when the view reopens.
OhsMcpClientqueries the configured stateless Streamable HTTP MCP endpoints concurrently through Obsidian's desktop HTTP API, avoiding browser CORS requirements for loopback services. Discovered tool names and input schemas are cached per endpoint for five minutes, avoiding repeatedtools/listcalls while preserving a short compatibility refresh window.- Hybrid Chat sends only the current, directive-free question to OHS. Earlier chat turns remain available to the generation model but never alter retrieval. OHS performs hybrid retrieval and applies its native cross-encoder reranker when enabled.
- Explicit YAML directives map to OHS frontmatter filters:
@property(status=todo)includes an exact value,@property(status!=done)excludes one, and@property(publication_date)requests a property without filtering. - Conversation history remains available only to the generation model; retrieval never uses earlier turns.
FederatedRetrievermerges per-vault ranks without comparing raw cross-vault scores, and a diversity pass gives each healthy vault a chance to contribute. When the opt-in related-note setting is enabled and the question explicitly asks about relationships, the best direct result from each healthy vault is expanded through one vault-local hop of links and backlinks; at most two linked candidates per anchor are promoted. - Only the globally selected note paths are sent to OHS
read. If a selected path is missing or a vault read fails, the retriever moves down the ranked candidates until the requested number of readable notes is filled or candidates are exhausted. Unavailable endpoints remain visible partial failures. ContextPackerapplies per-note and total character limits and centers long-note excerpts around the matching OHS snippet instead of always taking the beginning. Each source is namespaced asvault_id::vault/relative/path.md. Only explicitly requested YAML properties from the currently open vault are appended; current OHSreadresponses do not expose arbitrary cross-vault frontmatter.- A fresh local/UTC timestamp and optional custom instructions are added to the protected grounding prompt. No datetime MCP call or model tool selection is required.
OpenAiCompatibleChatClientsends the current question and a bounded recent transcript to/v1/chat/completions, then streams SSE output. A new chat starts with empty conversational memory. Retrieval never depends on model tool-calling.- The answer's source markers are matched to the notes included in the model's context. Cited notes appear in an expanded Cited sources list, while unused context notes appear under collapsed Other retrieved sources. Citations in the current vault open through the Obsidian workspace API. Other-vault citations use validated, percent-encoded
obsidian://open?vault=…&file=…URIs.
The source is split into the OHS client, federated retriever, rank fusion, context packer, OpenAI-compatible client, citation mapper, settings, and session/UI layers.
Chat window in the right-hand sidebar, source opens in the main editor:
Reading citations
Hybrid Chat accepts source markers such as [S1], [S1, S2], [1], and 【S1】, including harmless spacing or invisible characters inside the brackets. It renumbers cited notes in their first-mentioned order, starting at [1]. Markers for unknown sources are shown as unverified. Matching a marker to a retrieved note does not independently verify that the note supports the claim.
The source lists contain notes whose excerpts fit into the context sent to the chat provider. Notes found by search but excluded by the context limit are not shown as answer sources. If the answer cites none of the included notes, Hybrid Chat says so and keeps the notes under Other retrieved sources. Copying a message uses the displayed citation numbers; copying a whole chat as Markdown also includes the source grouping.
Some providers insert narrow no-break spaces between ordinary words, which can prevent line wrapping in the sidebar. Hybrid Chat converts those spaces to regular spaces for display and copying; the original provider response remains stored in the session.
Privacy boundaries
- OHS endpoints may be local or remote. Every selected endpoint receives only the current, directive-free question; earlier user questions and assistant answers are not included in OHS queries.
- The configured chat provider receives the packed source text and up to 12 recent non-empty chat messages within a 12,000-character transcript budget. The current question is always retained. A loopback provider can keep generation local; a cloud provider sends that data to the provider.
- Remote chat endpoints must use HTTPS. Plain HTTP is accepted only for loopback hosts.
- API keys are stored by Obsidian
SecretStoragethroughSecretComponent. Plugin data contains only secret identifiers. - Chat sessions and retrieved source text are stored in this plugin’s
data.jsonso sessions survive restarts. If the vault configuration is synchronized, that plugin data may synchronize too. - Source text is marked as untrusted context in the system prompt, but prompt injection remains a model-level risk.
- This MVP has no note mutation, Redis, embeddings, Graphify, Docling, OCR, direct SQLite access, or OHS reindex/status calls.
Configuration
In Obsidian settings:
- Register one OHS Streamable HTTP MCP endpoint per vault, including a stable ID, display name, exact Obsidian vault name, and enabled/default-selection state.
- Set a per-endpoint request timeout. The default is 60 seconds for each OHS direct search, related-note traversal, or read; a client timeout stops Hybrid Chat from waiting but cannot cancel synchronous database work already running inside OHS.
- Configure one or more OpenAI-compatible profiles, choose an active profile, and select or create an API-key secret. Use Refresh to load model IDs from the provider's
/modelsendpoint and choose one, or enter a model ID manually when discovery is unavailable. - Optionally customize language, tone, role, or answer structure. Grounding/citation rules remain enforced separately.
- Keep the local current-date/time injection enabled when relative dates such as “today” or “last week” matter.
- OHS reranking is enabled by default. Disable it when lower latency matters more than precision. The reranker model is configured on each OHS server, not in this plugin.
- Optionally enable one-hop related notes for relationship-oriented questions. It is disabled by default and never runs for ordinary factual searches.
- Use explicit YAML directives in chat prompts:
@property(field=value)filters by an exact value,@property(field!=value)excludes it, and@property(field)adds that current-vault value to answer context without filtering. - Adjust per-vault search, global note-read, and context limits.
Improving retrieval
Write the current question as a descriptive phrase or sentence that includes distinctive names, topics, dates, document types, and relationships. Hybrid Chat sends that question directly to OHS semantic/hybrid search; it does not reduce it to individual keywords. Because earlier turns are not part of retrieval, restate the subject in follow-up questions. When possible, search only the likely vault, use @property(...) filters to narrow the corpus, and enable one-hop related notes for questions explicitly about links or relationships.
For a balanced increase in recall, start with 24 results per vault, 10 notes read, 32,000 total context characters, and 4,000–5,000 characters per note, with OHS reranking enabled. For unusually difficult, high-recall searches, try 32 results per vault, 12 notes read, 48,000 total context characters, and 4,000 characters per note. The settings UI allows up to 50 results per vault and 20 notes read, but maximizing both can add latency and weakly related evidence.
Raise results per vault and notes read together: extra candidates normally cannot affect the answer unless enough globally ranked notes are subsequently read. Also raise the total context budget when reading more notes, or later sources may not fit in the provider prompt. Larger values are not always better; reduce them again when answers become slower or less focused.
The current public OHS contract exposes search and batch read over stateless Streamable HTTP. Tool prefixes and input schemas are discovered from each endpoint.
Streamable HTTP remains the default because multiple clients can share one long-lived OHS indexer and model cache. STDIO is not currently exposed as a Hybrid Chat transport; a future local-only mode would need to own a persistent child process and make its indexing, memory, logging, and cancellation lifecycle explicit.
Development
npm install
npm run lint
npm run typecheck
npm test
npm run build
npm run check
For manual testing, copy manifest.json, main.js, and styles.css into <vault>/.obsidian/plugins/hybrid-chat/, enable the plugin, configure OHS and a chat profile, then use the ribbon icon or Hybrid Chat: Open chat command.
Releases
main.js is generated and intentionally not committed. Run npm run package-release to execute every quality gate, validate that package.json, manifest.json, and versions.json agree, and create these files under dist/release/:
main.jsmanifest.jsonstyles.csschecksums.sha256
Pushing a tag that exactly matches manifest.json (for example, 0.1.8) runs the release workflow against the tagged source and generates build-provenance attestations from checksums.sha256. The workflow creates a draft release if needed, then uploads and verifies the three supported Obsidian assets: main.js, manifest.json, and styles.css. If a release for that tag already exists, it fills missing assets and checks existing assets against the tagged build without changing its notes or publication state. A new draft prepends the matching changelog entry to GitHub's generated notes. Review the draft before publishing. To repair an existing release after updating the workflow, run Release Obsidian plugin manually from GitHub Actions and enter the existing tag; the workflow checks out that tag before building.