Advanced Audio Recorder

by akhmialeuski
5
4
3
2
1
Score: 58/100

Description

Advanced audio recording plugin for Obsidian — multi-track capture, pause/resume, format conversion (WAV, MP3, FLAC, OGG, WebM, MP4), configurable save location, and built-in diagnostics.

Reviews

No reviews yet.

Stats

16
stars
3,371
downloads
2
forks
139
days
8
days
8
days
109
total PRs
0
open PRs
0
closed PRs
109
merged PRs
8
total issues
0
open issues
8
closed issues
780
commits

Latest Version

8 days ago

Changelog

This release lets you repair a diarized transcript line by line, right in the note. Split selection into another speaker hands the selected words of one line to another speaker, and Merge selected lines into one joins consecutive lines of one turn back together. Both edits rewrite the note and every transcript file of the recording from the same segments. The Transcribe audio dialog now fits on one screen, with its secondary options in a collapsed More options block, and a quick note can be transcribed by an engine of its own. Existing recordings, stored settings, and the recorder are unaffected.

New: Splitting a transcript line between speakers

When two people talk over each other, or the engine guesses wrong, one transcript line holds words from both of them. Select those words inside the line, right-click, and choose Split selection into another speaker, which is also a palette command.

  • Spoken by picks who the selection belongs to: any speaker the transcript shows, a participant of the recording nobody is named after yet, or a new Speaker N.
  • Time span starts from the word timings when the recording's JSON transcript has them. A bar under the fields covers the line with a handle at each end of the selection, so the boundary is dragged or typed, and the play button plays the span with a playhead running along the bar.
  • A selection in the middle of a line splits it in three, and Rest of the line spoken by picks the speaker of the words after it.
  • The note line is replaced through the editor, so Ctrl/Cmd+Z undoes it. The JSON transcript's segments are cut at the entered times, each part keeping its word timings, and the JSON, SRT, WebVTT, and TXT files the run wrote are rewritten from them. A new speaker joins the recording's roster, so Rename speakers plays and names it.

New: Merging lines into one

Diarization also cuts one person's turn into several lines when it hears a pause as a change of speaker. Select the consecutive lines and choose Merge selected lines into one from the editor menu or the palette. The dialog shows the merged text, starts Spoken by on the first line's speaker, and plays the whole passage, so you can check it really is one voice. Two lines are merged at once, and three or more are confirmed first. The note line and the transcript files are written the same way as for a split.

Without a JSON transcript, or when its segments do not match the lines (after a hand edit, for example), both edits change only the note, and the dialog says so before anything is written.

New: More options in the Transcribe audio dialog

The dialog had grown tall enough that Transcribe needed a scroll to reach. Language, the participant profile, and the destination stay up front beside the cost estimate, and every other option moves into a collapsed More options block. The block stays open while you change options, and it opens by itself when the stored engine cannot run on this device, so the reason Transcribe is disabled is in view.

Fixed

  • A quick note was always transcribed by the engine recordings use, because the only engine row on the Quick notes page picked the LLM that rewrites the text, and its name read as the engine for the whole dictation. The page now has a Quick note transcription engine row, and the dictation, its size limits, and its recorded cost all follow it. The old row is renamed Quick note rewrite engine.

Internal

  • The note outputs a transcription writes now record their timestamp template in the recording's sidecar, so a line is read back with the templates it was written with even after the settings change.
  • The line model the split and the merge share lives in its own module, and the two dialogs share one harness in the integration suite.
  • The suite is 6483 tests across 261 suites.

Compatibility

Requires Obsidian 1.6.6+, unchanged. This release is backward compatible: existing recordings, stored settings, the recorder, and the players are unaffected. Settings written by an earlier version carry no quick note transcription engine, which then starts on the engine recordings use, so a dictation is transcribed exactly as before until the row is changed. A transcript written by an earlier version has no stored timestamp template, and its lines are read with the current setting.

Full Changelog: https://github.com/akhmialeuski/advanced-audio-recorder/compare/2.3.3...2.3.4

README file from

Github

Advanced Audio Recorder for Obsidian

Advanced Audio Recorder lets you record, play, clean up, and transcribe audio directly inside Obsidian.

Use it for voice notes, meetings, interviews, lectures, and dictation. Record audio with one click, play it back in an enhanced waveform player, add bookmarks and chapters, split or convert files, and clean up noisy recordings when needed.

The plugin also includes built-in transcription. You can transcribe recordings or existing audio files in your vault using the OpenAI-compatible Whisper API (OpenAI, Groq, and other hosts), Deepgram, Google Gemini, Mistral Voxtral, or a local offline whisper.cpp model. Transcripts can include speaker labels, clickable timestamps, and an optional AI summary saved next to your note.

All audio and generated files stay in your vault. API keys are stored locally and are never sent anywhere except to the providers you choose, which are the transcription engine a run uses and, when post-processing is on, the LLM vendor it calls.

Works on desktop and mobile (iOS and Android). Some features are desktop-only (multi-track recording, local whisper.cpp transcription, input device selection); see Mobile support for the platform differences. Requires Obsidian 1.6.6 or newer. MIT licensed.

The full documentation lives in the docs folder, and the same link is built into the plugin's settings tab.

The enhanced audio player showing the waveform, playback controls, and a list of bookmarks and chapters.

Features

  • Recording: one-click capture with pause, resume, live status feedback, automatic splitting of long sessions, and crash recovery. Record in stereo or mono - including keeping just one channel of a dual-input audio interface.
  • System audio: record this computer's own output beside the microphone with one switch, so the remote participants of a call reach the recording (Windows).
  • Multi-track recording: capture up to eight input devices at once for multi-microphone interviews.
  • Enhanced audio player: inline waveform with adjustable speed, skip, loop, volume, a voice-boost toggle that applies the cleanup chain to the recording as it plays, per-recording bookmarks, chapters, and clickable timestamp links.
  • Transcription: OpenAI-compatible Whisper API, Deepgram, Google Gemini, Mistral Voxtral, or fully offline whisper.cpp, with speaker diarization and JSON, SRT, WebVTT, or plain-text output.
  • LLM post-processing: optionally clean up or summarize any transcript with OpenAI, Anthropic, Gemini, Mistral, or DeepSeek.
  • Quick notes: dictate into the note at the cursor from a ribbon button, with the audio never saved and an optional LLM rewrite from a profile of your own.
  • Audio cleanup: high-pass filter, noise gate, and loudness leveling, written to a fresh copy.
  • Formats and file operations: convert between WAV, WebM, OGG, MP3, MP4, M4A, AAC, and FLAC (with an optional mono downmix), and split long files from the right-click menu.
  • Command line (desktop, Obsidian 1.12.2+): start or stop a recording, ask what the recorder is doing, or transcribe a vault file from a terminal.

Recording is fast and forgiving. Start and stop from the ribbon or a command, follow live feedback in the status bar, and pause or resume without losing anything. Capture up to eight input devices at once for multi-microphone interviews, let long sessions split into fixed-length parts automatically, and recover the audio on the next launch if Obsidian closes mid-recording. See Recording and Multi-track recording.

The recording status bar showing the Recording label, control buttons, elapsed time, file size, and input level meter.

The enhanced player turns playback into a first-class part of your notes. Recordings embed as a player with a clickable waveform, adjustable speed, skip, loop, and volume. Add per-recording bookmarks and chapters, move between them, and copy timestamp links that jump straight to a moment in the audio. See the enhanced audio player.

Transcription turns any recording, or any audio file already in your vault, into text with the engine that fits your needs: the OpenAI-compatible Whisper API (Groq and other compatible hosts included), Deepgram, Google Gemini, Mistral Voxtral, or a fully offline whisper.cpp model that never touches the network. Deepgram, Gemini, and Voxtral add automatic speaker diarization, so meetings and interviews come back labelled by speaker. You decide where the transcript goes and in which format (JSON, SRT, WebVTT, or plain text), with timestamps you can click to jump the player to the right moment. An optional pass through an LLM (OpenAI, Anthropic, Gemini, or Mistral) cleans up the wording or condenses the transcript into key points and action items. See Transcription and LLM post-processing.

The Transcribe audio dialog with its per-run rows for engine, language, speaker diarization, the participant and dictionary profiles, the two-pass mode, the destination and file format, LLM post-processing and chapter generation

Everything else keeps your audio tidy. Convert recordings between WAV, WebM, OGG, MP3, MP4, M4A, AAC, and FLAC, and split long files into parts straight from the right-click menu; the formats you can record in are the subset your platform supports. Clean up noisy audio on demand with a high-pass filter, noise gate, and loudness leveling, always written to a fresh copy so the original is left untouched. See Formats, File operations, and Audio cleanup.

Installation

  1. In Obsidian, open Settings, go to Community plugins, and turn off Restricted mode if it is on.
  2. Click Browse, then search for "Advanced Audio Recorder".
  3. Click Install, then Enable.

Turn Obsidian's own Audio recorder off. Obsidian ships a core plugin of that name which puts its own microphone button in the same ribbon, and two recording buttons side by side are easy to confuse. Switch it off under Settings > Core plugins. To keep the core plugin and drop only its button, right-click an empty part of the ribbon strip instead and untick its entry. See Getting started.

The Audio recorder core plugin with its toggle switched off in Settings, Core plugins

Quick start

  1. Click the microphone icon in the left ribbon, or run "Start/stop recording" from the command palette.
  2. Speak, then click again to stop.
  3. The recording is saved to your vault and embedded in the active note.

No default hotkeys are assigned. Set your own in Settings, under Hotkeys.

Support

If this plugin saves you time, consider supporting its development: Buy Me A Coffee.

Troubleshooting and bug reports

If something is not working, start with the troubleshooting guide. When you report a problem, follow the bug reporting guide and open an issue on GitHub.

License

Released under the MIT License.