Alvaro's Techincal Dives

Classifying Standard Notes Tags Using Local AI Models

I have years of notes in Standard Notes: journal entries, project scratchpads, half-finished ideas. Almost none of them are tagged consistently, and I wanted a language model to read through them and suggest tags. The obstacle is that Standard Notes is end-to-end encrypted. The service stores ciphertext, and the decryption keys are derived from my password on my devices. Sending a decade of personal writing to a cloud API to get tagging suggestions would undo the one property I chose the service for.

So the constraint was: the model runs on my machine, and the plaintext never leaves it. That sounds like one decision, but it forced two harder problems that don't usually appear together. First, I needed to pull my own notes out of an encrypted vault in bulk, without the official client. Second, I needed a local model that stayed reliable across a long batch job. This post covers both, and the places where they fought each other.

The pipeline at a glance

The project is a set of small Node.js scripts in packages/snjs-get-all-notes, written in TypeScript and depending only on libsodium-wrappers-sumo, node-fetch, dotenv and prompt-sync. It does not use @standardnotes/snjs. Each stage writes files that the next stage reads, so any step can be inspected or rerun on its own:

Stage Entry point Reads Writes
1. Fetch and decrypt index.ts Sync API output/notes.json, output/tags.json, two summary .txt files
2. Clean and filter transform.ts output/*.json output/clean/*.json, then AI tagging into output/ai-tagged/
3. Push new tags and links src/sync-tags.ts output/ai-tagged/notes.json Writes to the vault
4. Deduplicate src/dedup-tags.ts output/tags.json Deletes duplicates in the vault

The protocol code lives in src/api.ts (HTTP, sync, writes), src/crypto.ts (key derivation and the V004 cipher), and src/decrypt.ts (the key hierarchy applied to notes and tags). Everything below follows that layering.

One thing to state up front: stage 1 writes decrypted notes to output/ as plain JSON and text. The privacy guarantee applies to the network, not to your disk. Anything in that directory is readable by anyone with access to the machine, and the password sits in a plaintext .env file. I'll come back to both in the limitations.

Part 1: Signing in without the official client

Standard Notes login is a two-request exchange, plus an optional MFA step. The first request, POST /v2/login-params, asks the server for the key derivation parameters for the account: a pw_nonce, a protocol version, and an identifier. The client must answer with a PKCE code_challenge, which the server later checks against the code_verifier sent in the second request.

The challenge is not the standard one. The standard is base64url(SHA256(verifier)) over the raw digest bytes. Standard Notes instead hashes to a hex string, then base64url-encodes the UTF-8 bytes of that string:

function computeCodeChallenge(codeVerifier: string): string {
  const hashHex = crypto.createHash('sha256').update(codeVerifier).digest('hex')
  return Buffer.from(hashHex, 'utf8').toString('base64url')
}

The code comment cites the web client's base64URLEncode(sha256(...)), where sha256 returns hex and the encoder encodes its characters. I only discovered this because a mismatched challenge produced an authentication failure rather than a helpful error, so the failure looked like a wrong password. If you implement this yourself, verify the challenge against a known-good login before debugging anything else.

The client rejects any protocol version other than 004 with an explicit error. Older accounts on 003 are not supported; that choice is deliberate and visible in the code.

Key derivation

Key derivation is where the security actually lives. The client takes the password and derives 64 bytes with Argon2id:

const saltHash = crypto.createHash('sha256').update(`${email}:${pwNonce}`).digest('hex')
const salt = hexToBytes(saltHash.slice(0, 32)) // first 128 bits

const derivedKey = sodium.crypto_pwhash(
  64, password, salt,
  5,                  // opslimit (time cost)
  64 * 1024 * 1024,   // memlimit: 64 MiB
  sodium.crypto_pwhash_ALG_ARGON2ID13,
  'hex',
)

const masterKey      = derivedKey.slice(0, 64)   // bytes 0-31
const serverPassword = derivedKey.slice(64, 128) // bytes 32-63

Three details matter here.

The salt is not random. It's the first 16 bytes of SHA-256 over email:pw_nonce. The nonce is per-account and comes from the server, so the salt is stable for an account and does not need to be stored locally. It also means the email address participates in the salt, so the same password on two accounts produces different keys.

The 64-byte output is split. The first half is the master key, which never leaves the client. The second half is the serverPassword, which is what the login request sends. The server stores a hash of that half and can verify it, but cannot derive the master key from it. This is the reason login doesn't reveal the key that decrypts your vault.

The memory and time parameters are a cost you pay on every sign-in. Argon2id at 64 MiB with opslimit 5 is designed to be slow for an attacker running many guesses, and it is slow for the legitimate client too. The script waits on it once per run, which is acceptable for a batch job and would be a noticeable delay in an interactive UI.

Session tokens

The login response returns a session, and the choice of API version determines its shape. Standard Notes has two session formats. The newer one (api=20240226) is cookie-based: the usable secret arrives only in a Set-Cookie header, and the JSON body carries an identifier. The older one (api=20200115) returns the full token as 1:<uuid>:<secret> in the body, usable as a Bearer token from a script. The client pins API_VERSION = '20200115' for every request because it is the only version where a plain Node.js process gets a token it can use directly.

This is a dependency on an older API surface, and it may stop working. I'd treat it as a known risk rather than a stable interface.

If the account has MFA, the server responds with an mfa-required error that carries an mfa_key. The client prompts for the OTP and resubmits the same request with the code under that key. The code is taken from an environment variable if present, so the flow can run unattended after one manual entry.

Part 2: The V004 ciphertext

Every encrypted field in the sync payload uses the same string format:

004:<nonce_hex>:<ciphertext_base64>:<aad_base64>

The cipher is XChaCha20-Poly1305 (libsodium's crypto_aead_xchacha20poly1305_ietf). It is an AEAD scheme, so decryption verifies an authentication tag over both the ciphertext and the associated data. The nonce is 24 bytes, which is why random nonces are safe here: the space is large enough that collisions are not a practical concern.

The AAD is the subtle part. The stored aad field is the base64 of a JSON object with sorted keys and nullish values removed. That object identifies the item and the protocol version. At decryption time the client does not trust the stored blob alone. It rebuilds the AAD from live fields:

function buildCanonicalAAD(item: RawItem, encryptedString: string, storedAAD: Record<string, unknown>): string {
  const version = encryptedString.substring(0, 3) // "004", not sent by the server
  const merged = { ...storedAAD, u: item.uuid, v: version }
  if (item.key_system_identifier) merged.ksi = item.key_system_identifier
  if (item.shared_vault_uuid) merged.svu = item.shared_vault_uuid
  return Buffer.from(sortedJson(merged), 'utf8').toString('base64')
}

Several things follow from this design:

  • The item's UUID is bound into the ciphertext. A ciphertext copied from one note and placed under another note's UUID fails authentication. Encrypted content cannot be silently moved between items.
  • The protocol version is bound too. The version comes from the ciphertext's own prefix, so a 003 string cannot be passed off as 004.
  • Vault membership is bound. The ksi and svu fields tie an item to its key system and shared vault, which means moving an item between vaults invalidates it.
  • The serialization has to be exact. sortedJson sorts keys and drops null and undefined. A single difference in key order or whitespace yields a different AAD, and XChaCha20-Poly1305 then fails with an authentication error that doesn't say which part was wrong. Most of my debugging time went into this function.

Decryption is layered. Standard Notes uses a three-level hierarchy, and each level is its own AEAD operation:

masterKey
  └─ decrypts enc_item_key of SN|ItemsKey  → contentKey
       └─ decrypts content of SN|ItemsKey  → { itemsKey }
            └─ itemsKey decrypts enc_item_key of Note/Tag → contentKey
                 └─ contentKey decrypts content of Note/Tag → plaintext JSON

The reason for the hierarchy is rotation and isolation. Each item gets its own content key, so compromising one item key does not expose the others. The vault's ItemsKey is itself encrypted under the master key, so changing the password only requires re-encrypting that small item. The notes themselves don't need re-encryption on a password change.

Two smaller details in src/decrypt.ts:

  • Items that fail with Unsupported protocol version are skipped silently. These are legacy items that the script cannot read. Other decryption failures are logged. The script therefore counts only what it could read, and a silent skip looks like a missing note, so the logs are the place to check.
  • Timestamps come in two forms. The updated_at string is millisecond-precision, and updated_at_timestamp is microseconds. The code prefers the microsecond value when present. The precision matters later, because the server compares timestamps exactly when it checks writes.

Part 3: Syncing everything

The sync endpoint is POST /v1/items. Pulling everything means sending an empty items array with limit: 300 and following cursor_token until it's null. The sync_token returned on each page is the server's marker for the client's state, and it must be sent back on writes so the server knows which version the client is building on.

do {
  const body = { api: API_VERSION, items: [], limit: 300 }
  if (syncToken)   body.sync_token   = syncToken
  if (cursorToken) body.cursor_token = cursorToken
  // POST /v1/items, collect retrieved_items, drop deleted ones
  cursorToken = syncData.cursor_token ?? null
} while (cursorToken)

Deleted items come back in the sync response, so the client filters them out locally. Standard Notes deletion is always soft. The server sets deleted: true and clears the content, and the client removes the item from its view. The script keeps the sync token from the last page so that later writes build on the freshest server state.

Part 4: Preparing notes for a model

The raw note text field is not plain text. Standard Notes stores rich text as Lexical JSON, spreadsheets as their own JSON structure, and plain notes as plain strings. src/transform.ts normalizes all three:

  • Lexical: it walks the node tree, collects text leaves, and joins block-level nodes such as paragraphs, headings, list items, quotes and code with newlines. Excessive blank lines are collapsed.
  • Spreadsheet: it flattens each sheet into tab-separated rows, aligning cells by their column index so sparse rows keep their columns, and prefixes each sheet with [Sheet: name].
  • Anything else: an unparseable JSON-looking string is returned as-is, and an unknown JSON structure is re-stringified compactly. That fallback is a guess, not a guarantee, and it's worth checking against a real note of each type before trusting the output.

There's a filter in transform.ts that is easy to miss: only notes with three or fewer tags are sent to the model. The reasoning is that notes already carrying several tags are presumably categorized, and the goal was to suggest tags for the rest. This makes the AI stage a gap filler rather than a re-categorizer. If you want the model to reconsider well-tagged notes, that filter is the first thing to change.

Part 5: The local model stage

The server and the contract

src/ollama.ts talks to LocalAI through its OpenAI-compatible /v1/chat/completions endpoint. The defaults are http://localhost:8080 and the gemma-4-e4b-it model, both overridable through environment variables. The file was written to target Ollama and later pointed at LocalAI

The output contract is one JSON array of { "uuid", "tags" } objects, with 1-3 word lowercase tags, 3-5 per note, and no markdown fences. Because local models don't follow formatting instructions perfectly, the parser is defensive. It tries JSON.parse on the whole response, falls back to extracting the first [...] block with a regex, and unwraps a single-key object such as { "notes": [...] }. Entries missing a UUID or a tags array are dropped silently, and a batch that yields nothing usable is logged and skipped. Each note's tags are then normalized with trimming and deduplication.

Gemma does not accept a system role in this setup, so the instructions are prepended to the user message. This is a small thing, but it's the kind of backend-specific detail that breaks the first time you change models.

One note per request

The batch size is a constant: BATCH_SIZE = 1. The code comment says this keeps inference under LocalAI's five-minute busy watchdog. A prompt that included the full vault tag list, roughly 97 entries at the time, made the model produce a chain of reasoning around 700 tokens long before it wrote an answer. Multiplied by a long prompt and a slow local model, that is how a request ran past the watchdog and returned a 500.

The fix was to remove the tag list from the prompt entirely and match against existing tags in code afterwards. tagNotesWithAI lowercases both sides and splits each result into existingTagsApplied and newTagsSuggested. The model's freedom to invent tags is therefore not a problem to be prompted away. It becomes an output to classify.

This trades throughput for reliability. Batching several notes per call would be faster if the model could handle it, but the watchdog is per request, so one slow batch fails all its notes together. A single note fails alone and is retried on the next run.

OLLAMA_CONCURRENCY (default 3) controls how many single-note requests are in flight. Concurrency and the watchdog interact: three concurrent inferences on one local model compete for the same hardware, so each one takes longer. I did not measure where that tradeoff sits on my machine, and the right value depends on the hardware. Setting it to 1 gives strictly sequential behavior at the cost of wall-clock time.

Context truncation

The model sees the title, the currentTags, and the first 200 characters of the body. That limit is in the code as slice(0, 200), and it's a deliberate tradeoff: longer prompts increase both latency and the chance of watchdog timeouts. The cost is that tags depending on content past the first 200 characters are invisible to the model. For a long journal entry, the opening lines may not reflect the topic. A better design would summarize or chunk longer notes, but that is more inference per note, which runs against the same budget.

Sanitizing text for the backend

Before the text reaches the model, it passes through sanitize:

  1. Characters above U+FFFF, such as emoji and rare symbols, are removed.
  2. Non-breaking and other irregular spaces become plain spaces.
  3. The text is NFD-normalized and combining marks are stripped, so ã becomes a and ç becomes c.
  4. Anything else outside printable ASCII is dropped.

My notes contain Portuguese, and stripping diacritics changes words: informação becomes informacao. The model handles the ASCII form adequately for tagging, but the stored tags come from the model's output, not the original text. A tag like informacao can end up in the vault if the model produces it. Anyone reading the post should know that the sanitization affects what the model reads, though not the notes themselves, which are untouched in the vault.

Progress tracking

A tagging run over hundreds of notes takes long enough to be interrupted, and a crash should not cost completed work. tagNotesWithAI writes three things:

  • progress.json records one BatchRecord per batch with a status of pending, running, done or failed, plus timestamps, the note UUIDs in the batch, the error message, and the number of tags assigned.
  • notes.json is merged after each batch completes. mergeAndSaveNotes keys by UUID, so a re-run that reprocesses a note overwrites it rather than duplicating it.
  • Writes are serialized through a promise chain (withLock), so concurrent batches don't interleave their writes to the progress file.

On startup, a compatible progress file (same model, host, batch size and note count) is reused. Batches left as running by an earlier crash are reset to pending, since they never finished. A progress file from a different configuration is replaced, which is the safe default, because stale batches would be attached to different notes or a different model.

Notes that fail in every batch are returned untagged, with empty AI fields, so the output always covers every input note. The summary at the end lists failures with their error messages.

Part 6: Writing tags back to the vault

Reading the vault is one direction. Writing tags back means becoming a client in the other direction, which is more constrained, because the server is checking the integrity of everything you send.

Creating tags

createTags builds a new Tag item for each missing title. It generates a random UUID, creates a fresh 32-byte content key, encrypts the content ({ title, references: [] }) under that key, then encrypts the content key under the vault's active ItemsKey. Both encryptions use the new item's UUID in the AAD, so the item is bound to its identity from creation.

Creation sends all items in one sync call and then checks saved_items against the UUIDs it sent. Any item missing from the response is reported as unconfirmed rather than assumed to exist, which matters because the write can partially succeed.

Linking tags to notes

A Standard Notes tag holds its own list of references, so linking a note to a tag means editing the tag, not the note. linkTagsToNotes decrypts each tag's current content, takes the non-note references as they are, sets the note references to the list passed in, re-encrypts with a fresh content key, and syncs the batch.

The function's documentation says the list is "the full set of note UUIDs that should be referenced by this tag after the update." That is what the code does. It replaces the note references rather than adding to them. This is the correct behavior for deduplication, where the caller passes the union of all duplicates' notes. It is not automatically correct for the tagging step, where the caller only knows the notes the model tagged. See the note in the limitations.

Each re-encryption uses a new content key. The rewritten tag content is therefore new ciphertext under a new key, rather than the old key reused with a different message, which would be a much more serious mistake.

Timestamps and conflicts

The server checks each write against its own record. The code refers to this as a TimeDifferenceFilter, and the failure mode is concrete: if the updated_at in a write doesn't match the server's, the item comes back as a sync_conflict. Stamping every write with new Date() causes every write to conflict, which is why the code copies timestamps from the item it fetched and sends sync_token with every batch.

The retry path is what made the writes reliable. When a conflict comes back, the server includes its authoritative server_item. The client takes that item's timestamps, re-encrypts the same desired content, and retries once after fetching a fresh sync token. The content is unchanged. Only the timestamps move forward to the server's version. Conflict types other than sync_conflict, such as uuid_conflict, are logged and not retried.

Pacing

Writes go out in batches of 20 with a two-second pause between them. The code comment attributes this to a five-minute bandwidth cap on the server. A 429, or a transient 502, 503 or 504 response, triggers a retry with exponential backoff: 30 seconds, doubling each attempt, for up to eight attempts.

Part 7: Deduplication

Duplicate tags happen when a run creates a title that already exists in a form the client didn't recognize, or when the same title is created twice. src/dedup-tags.ts plans the cleanup from output/tags.json, which is a snapshot from the last fetch:

  1. Group tags by lowercased, trimmed title.
  2. Within each group, collapse repeated UUIDs, keeping the entry with more notes.
  3. Pick the canonical tag: the one with the most notes, with the oldest updatedAt as the tiebreaker.
  4. Compute the union of note UUIDs across every copy in the group.
  5. If the canonical tag doesn't already reference the whole union, relink it to the union.
  6. Only after every relink succeeds, soft-delete the non-canonical copies.

The ordering is the important part. The script aborts before deleting anything if any relink fails. That makes the operation safe to rerun: a failed relink leaves every duplicate in place, and a subsequent run recomputes the plan.

The plan uses the snapshot from tags.json, not a live read. If the vault changed after the snapshot was taken, the plan may be stale. The script re-syncs before writing, but it relinks using the snapshot's note lists. A live re-plan would be safer, and the next version should build the plan from fresh data.

Limitations and what I'd change

Plaintext on disk. The run writes decrypted notes to output/, and the password is read from .env. Both are unencrypted at rest. The network path is protected by the protocol, but the local artifacts are only as safe as the machine and its file permissions.

Stale snapshots. Deduplication plans from a snapshot. The tagging step decrypts the live vault before writing, which is the better pattern, and the same approach should be applied to deduplication.

Model-dependent behavior. The tag count, the sanitization and the watchdog interaction were all tuned against one model on one machine. Changing the model or the hardware can change the right batch size and concurrency.

What I'd take from this

The crypto was the part I expected to be hard, and it was: the AAD canonicalization and the PKCE quirk each cost real time. The part that consumed more effort, though, was the operational side. Making a local model behave predictably through a long batch job meant working around a watchdog, a backend that doesn't like some Unicode, and a prompt that became expensive as soon as I gave it more context. Making the vault writes safe meant understanding the server's conflict checks well enough to recover from them instead of fighting them.

"Keep it local" is a good default for personal data. It is also a commitment to owning every failure mode, because no vendor is standing behind either the protocol or the model. If you're considering a similar project, budget most of the time for the boring parts: resumable progress, conservative writes, and a way to see exactly what the model did before anything reaches the vault.