Monday, August 31, 2026

Claude Text Watermarks and EU Provenance: What Publishers Should Disclose

Claude text watermarks and EU provenance

On 14 August 2026 Anthropic published how it will watermark future Claude text. The short version: new models will mark outputs; a mark is a likelihood, not a confession; you still have to say when you used the model. The long version is Article 50 of the EU AI Act, a July 2026 Code of Practice, and a detection API that is not shipping yet.

This post is for people who publish with Claude, ship a product on the Claude API, or have to write a disclosure line that will survive legal review. Sources: How Claude’s text watermark works (14 August 2026) and How Claude marks AI-generated content.

Two clocks

The EU rule for providers serving its market is already on. Anthropic’s implementation is staggered.

When What actually happens
2 August 2026 EU marking obligation for AI providers serving the EU market. Anthropic’s cutoff for “new models mark at launch.”
July 2026 Anthropic and other major providers signed the EU Code of Practice on Transparency of AI-Generated Content (~190 signatories).
14 August 2026 Anthropic explained the method: SynthID-Text-style token-choice watermark; C2PA credentials on supported files.
Coming months Detection API “soon.” Older models (launched before 2 August 2026) get marking during a transition period.

If your stack is still Sonnet 4 / Opus 4.x, do not assume every paste is already watermarked. Anthropic is explicit that models launched on or after 2 August 2026 support marking at launch, and that pre-cutoff models are “in progress.” Independent testers in mid-August reported nulls on some current Claude names; treat that as “not yet,” not as “never.”

What the watermark is (and is not)

Claude still picks the next token from the same candidate list. On low-stakes choices (“overcast” vs “grey”), the source of the randomness is a key plus recent tokens instead of a plain RNG. Google DeepMind published the family as SynthID-Text in Nature in 2024. Anthropic says internal tests and DeepMind’s own thumbs-up study showed no practical quality hit, no extra tokens, no extra price, no hidden characters.

It does not:

  • Identify a user, org, or chat
  • Prove a human did not write the piece
  • Prove a different model wrote it (other vendors have other keys, other methods)
  • Survive a full rewrite
  • Work well on short samples, dense facts, or “fix only the grammar” edits
  • Sit heavily in code, where the next token is often forced

A translation Claude writes is fully watermarked, because Claude chose every word. A proofread of your draft may not register at all.

Files are a different channel. Supported images (PNG, JPG, SVG, and similar) get a C2PA signed content credential in metadata: Claude was involved. That is not woven into pixels. Strip metadata, screenshot, or re-save and it is gone. Anthropic will offer a drop-a-file checker; any C2PA-aware tool can read the label.

Marks apply worldwide on supported models, including API, Claude, Claude Code, Cowork, Tag, and cloud partners (AWS, Google Cloud, Microsoft Foundry). Provenance metadata may not exist on every platform.

What publishers should disclose

The watermark is the provider’s Article 50 marking duty. Your duty as a deployer or publisher is separate. Anthropic says: if you build on Claude, assess what Article 50 requires of your product. Do not wait for their detection API to write your user-facing sentence.

A workable 2026 policy for a blog, newsletter, or docs site:

  1. Say when Claude (or any model) drafted or substantially rewrote the piece. A machine-readable mark is not a substitute for a human-readable line. Readers and regulators will not run a detector on every URL.
  2. Do not claim “this is 100% human” just because a detector is quiet. Pre-cutoff models, short blurbs, heavy edits, and other vendors all produce silence.
  3. Do not claim “Claude wrote this” just because a detector fires. Anthropic’s own FAQ: a mark means Claude was likely involved — author, heavy editor, or translator. It is not a plagiarism verdict and not a quality score.
  4. Keep file provenance if you ship Claude-generated images. Prefer formats that retain C2PA. Do not treat a PNG screenshot as equivalent.
  5. If you are in the EU market or selling into it, write the disclosure into the CMS template, not into a one-off author bio. The Code of Practice is about systems, not hero posts.

For API products: the watermark is at the model, so it is in the text you receive. You still need your own UI disclosure if you present that text as your product’s output. You cannot “turn off” the token-choice mark on a supported model.

What not to do

Do not treat a forthcoming detection API as a cheating oracle for students or employees. Do not fire someone, reject a paper, or nuke a PR because a likelihood score moved. Do not strip C2PA and then claim the image is unmarked in a legal sense — the Act cares about the act of marking, not about whether you later flattened the file. Do not assume US-only traffic exempts you; Anthropic is applying marks globally because it cannot yet scope by region. Do not wait for “older models to catch up” before you write the disclosure sentence.

A simple map

Signal What it supports What it does not support
Claude text watermark (supported model) “Claude was likely involved in some of these tokens” Authorship, cheating, other vendors, short text
C2PA on a Claude file “This file was processed by Claude, and whether it was tampered with after signing” A screenshot, a re-encoded JPEG, a PDF print
Your published disclosure line What a reader and a regulator can actually use Nothing, if you skip it and hope the watermark is enough
Third-party “AI detectors” (style tells) A different, noisier guess Anything you would bet a job on

The EU rule is a provider marking rule plus deployer transparency. Anthropic is doing the first with SynthID-style text and C2PA files. You still own the sentence on the page. Write it now. Update it when the detection API actually exists.

This post is part of our operator notes on shipping AI. Follow the series at amtocbot.com.

Monday, August 24, 2026

AI as Infrastructure: Value Moves Up-Stack

For a few years the AI conversation was about who had the biggest model. That is the wrong altitude now. Models still matter, the way CPUs still matter — as a layer you buy, swap, and budget for. The margin is moving up the stack: into products, workflows, evaluation, and the data those products sit on.

This post is for people who ship software, buy inference, or have to explain to a board why "we use GPT/Claude/Grok" is not a strategy. It is not a market-sizing deck.

The stack, in one picture

At Davos 2026, NVIDIA's Jensen Huang described AI as a five-layer cake: energy, chips, computing infrastructure, models, and applications. Each layer has to be built and paid for. The application layer is the only one end users touch, and it is the reason the four layers underneath exist.

Two things follow.

First, the bottom of the cake is capital-intensive and crowded. Hyperscalers are pouring unprecedented capex into data centers, accelerators, and networking. That spend is real. It is also not where most product companies will differentiate.

Second, inference is now the production workload. Training still happens, but the day-to-day cost of AI is tokens out the door. Industry commentary through 2026 has inference taking a majority of AI compute, with agentic and reasoning workloads as the fastest-growing slice. If you run a product, you are in the inference business whether you meant to be or not.

Why models commoditize

A model is a capability. Capabilities leak.

Open weights, falling inference prices, and "good enough" alternatives mean the gap between the frontier API and the second-best option keeps shrinking for a large class of tasks. Routing, distillation, and small specialists eat the middle. The brand of the model still matters for a few flagship surfaces. It does not matter for most internal workflows.

That is the same pattern as cloud VMs, then containers, then managed databases. The undifferentiated layer gets cheaper and more interchangeable. Buyers stop paying a premium for "we have compute" and start paying for "this job is done."

What does not commoditize as fast:

- Proprietary data you can legally use in the loop

- Evaluation that matches the actual job (not a public leaderboard)

- Workflows that already live in the customer's day

- Distribution: the place the user already is

- Trust, audit, and on-call when the model is wrong

Those sit above the model.

Where the margin actually is

If the model is infrastructure, the product is the control plane.

Think of inference the way you think of a database. You do not advertise "we use Postgres." You advertise the workflow: the ticket that closes, the draft that ships, the claim that is coded, the incident that is triaged. The database is a line item. The workflow is the company.

Three practical consequences:

1. Switching cost lives in integration, not in the model card. Prompt libraries, tool schemas, eval sets, and human review queues are the lock-in. If those are thin, a competitor can swap your model next quarter.

2. Unit economics are tokens plus people. An agent that spends $0.04 of inference and $4 of human cleanup is not an agent product. It is a demo. Measure cost per completed job, not cost per million tokens.

3. Routing is a product decision. Different jobs want different models: cheap/fast for classification, stronger/slower for irreversible actions, local for data that cannot leave. The routing policy is yours. The vendors will all claim to be the only layer you need.

What to do this quarter

If you run a product or an internal platform:

- Treat the model API as a vendor, with a backup. Write a one-page "we can switch in 30 days" test and actually run it on one workflow.

- Put evals next to the feature, not in a slide. A frozen set of real tickets/emails/PRs beats a public benchmark.

- Own the workflow artifact: the ticket, the document, the PR, the claim. That is the up-stack asset.

- Budget inference as COGS, not as R&D theater. If you cannot say cost per successful task, you cannot say whether the feature should exist.

If you buy AI for a team:

- Ask "which job gets shorter?" not "which model is smartest?"

- Prefer tools that sit in the existing system of record. A new chat window is the down-stack move.

If you invest or advise:

- The crowded trade is GPUs and frontier brands. The quieter trade is the control plane: eval, routing, permissions, and industry workflow.

- Be suspicious of stories that stop at "we have access to a model." That is table stakes in 2026.

What not to do

Do not freeze the architecture on one vendor's chat API. Do not skip evals because the demo was impressive. Do not confuse "employees have Copilot seats" with "we captured workflow margin." Do not wait for the model layer to stabilize before you own the job — the model layer is supposed to keep moving. That is what infrastructure does.

A simple map

LayerWhat you buyWhere the margin is
Energy / chips / clustersCapex, cloud commitHyperscalers and hardware
ModelsAPI or weightsLabs, briefly; then price
Inference servingTokens, latency, regionUtilities, unless you own routing
Apps and workflowsCompleted jobsYou, if you own the loop

The punchline is not that models are worthless. It is that they are becoming plumbing. Plumbing has to be reliable, billed, and replaceable. The company that wins is the one whose product still works when the pipe is swapped.

This post is part of our operator notes on shipping AI. Follow the series at amtocbot.com.

How WebSockets Work — LearningTechBasics

LT LearningTechBasics @amtocbot

How WebSockets Work

Upgrading a one-shot request into a two-way live wire.

📅 2026-08-14⏱️ ~5 min read🏷️ Networking · WebDev

Plain HTTP is request-response: the client asks, the server answers, done. WebSockets keep the connection open so either side can send messages at any time — ideal for chat, live data, and games.

Legend — how to read this diagram

A · BPartiesthe two sides of the exchange
1–nOrdereach message, numbered in sequence
1 2 3Walkthroughnumbered steps below run in order

From HTTP to a live socket

  1. Upgrade request. The browser sends a normal HTTP request with an Upgrade: websocket header.
  2. 101 response. The server agrees and switches protocols on the same TCP connection.
  3. Full duplex. Both sides now send lightweight frames whenever they want.
  4. Close. Either side sends a close frame to end the conversation.

When to reach for them

Push, not poll. No repeated requests asking 'anything new?' — the server just tells you.

Low overhead. Frames are tiny compared to full HTTP requests.

Not always needed. For occasional updates, Server-Sent Events or polling may be simpler.

One-line mental model:

Start as HTTP, then upgrade the same connection into a persistent two-way channel.

Found this useful? Three ways to go deeper — start free.
Free · Weekly One tech idea, in your inbox

The same clear explainers, delivered weekly. No spam, unsubscribe anytime.

Subscribe free →
Guide · $39 The Open-Source AI Stack

120+ page production guide: local LLMs, fine-tuning, RAG, and domain-specific AI.

Get the guide →
Consulting Need this built properly?

AI implementation, fine-tuning and RAG done right. Free 30-minute strategy call.

Book a call →

Sunday, August 23, 2026

Let's Encrypt's Post-Quantum TLS Timeline: What Site Owners Change, and When

On 3 June 2026, Let's Encrypt published its plan for a post-quantum-safe Web PKI. The short version: your current certificates do not change today. The long version is a timeline, a new certificate design, and one server setting that actually matters this year.

This post is for people who run websites, terminate TLS, or maintain ACME clients — not for cryptographers. Source: A Post-Quantum Future for Let's Encrypt (Andrew Gabbitas, 3 June 2026).

Two different post-quantum problems

TLS does two jobs. Encryption hides the bytes. Authentication proves you reached the right server.

Encryption is the urgent one. An attacker who records traffic today can try to decrypt it later, once a cryptographically relevant quantum computer exists. That is the "harvest now, decrypt later" problem. The fix is already shipping: hybrid post-quantum key exchange, typically X25519MLKEM768 (classic X25519 plus NIST's ML-KEM). Major browsers and operating systems already support it. If your server does too, those connections are protected against future decryption of recorded sessions.

Authentication is slower to migrate. A quantum computer has to forge a signature in real time, not retroactively. Let's Encrypt still treats it as work that has to start now, because root programs, libraries, and ACME clients take years to move, and because governments have put dates on the wall: NSA CNSA 2.0 aims national-security systems at 2030–2035; NIST's draft guidance would deprecate RSA-2048 and P-256 after 2030 and disallow them after 2035; the EU roadmap targets high-risk systems by the end of 2030 and broad migration by 2035. Google and Cloudflare have both said they will migrate their own services by 2029.

Why not just put ML-DSA on every certificate?

The obvious next step is to replace RSA/ECDSA signatures with ML-DSA, the NIST-standardized post-quantum signature scheme. Size kills that as a default for the public web.

ML-DSA-44 signatures are about 2,420 bytes. RSA-2048 signatures are 256 bytes; ECDSA P-256 signatures are 64 bytes. Public keys grow as well. A typical Web PKI handshake today carries five signatures and two public keys. Swap those for ML-DSA and one handshake can exceed 10 kilobytes. Cloudflare's measurements show a meaningful share of real-world TLS connections fail at that size; the rest get slower. Let's Encrypt will not flip that on as the default.

They are tracking ML-DSA in X.509 (RFC 9881) and in TLS, and Go 1.27 adding ML-DSA to the standard library. Those pieces still matter. They are not the issuance path Let's Encrypt is betting on for web scale.

Merkle Tree Certificates instead

The path they announced is Merkle Tree Certificates (MTCs).

A conventional CA signs each certificate individually. An MTC CA issues certificates in batches. One signature covers the batch. Browsers keep a separately updated "landmark" of those batch signatures. In the common case, the TLS handshake then carries one signature, one public key, and one inclusion proof — smaller than today's handshake, even with post-quantum algorithms. If a client's landmark is stale, a larger "standalone" form is the fallback.

Transparency is not bolted on after issuance. The certificate only exists as a leaf in a published Merkle tree. Let's Encrypt has run Certificate Transparency logs (append-only Merkle trees) in production since 2019, so the data structure is not new to them.

Chrome has said MTCs are its preferred path for post-quantum certificates on the public web. Cloudflare and Chrome are already running a feasibility experiment against live traffic. The IETF PLANTS working group is standardizing the design.

The dates

Let's Encrypt's own targets, as of the 3 June 2026 post:

- Late 2026 — staging environment that issues MTCs - 2027 — production-ready environment

That is infrastructure readiness, not "every site on earth is post-quantum authenticated." Browsers, root programs, libraries, and ACME clients still have to land support. Let's Encrypt is in the PLANTS and ACME working groups while those standards settle.

What you should do now

If you operate a website or TLS terminator

1. Keep renewing Let's Encrypt certificates the way you do today. Issuance and renewal do not change.

2. Turn on hybrid post-quantum key exchange: X25519MLKEM768. This is the highest-leverage change in 2026. It does not require a new certificate. It does protect recorded traffic against later decryption.

3. Do not wait for post-quantum certificates before doing (2). Encryption and authentication are on different clocks.

If you maintain an ACME client or a certificate pipeline

Start tracking the PLANTS working group and the mtcs@chromium.org list. Some of the coming issuance changes will need client-side support. Clients that show up ready when staging opens will matter more than another blog post about algorithms.

If you are a security lead writing a 2026–2027 plan

Put "enable hybrid KEX on every public TLS terminator" in this quarter. Put "evaluate MTC / ACME client support" on the 2027 calendar, tied to Let's Encrypt's production target — not to a panic date in 2026. Post-quantum certificates from Let's Encrypt are promised to arrive the usual way: free, automated, ACME.

What not to do

Do not rotate to a paid CA just to "get PQ certs" in 2026. Public post-quantum certificates are not the bottleneck this year; handshake size and ecosystem readiness are. Do not disable classic certificates early. Do not treat a staging CA as production. Do not confuse "my CDN already does PQ key exchange" with "my origin certificate is post-quantum signed" — those are different layers.

A simple timeline

WhenWhat actually changes for you
Now (2026)Enable X25519MLKEM768 on servers. Certificates stay as they are.
Late 2026Let's Encrypt staging MTCs. Watch if you run ACME software. Ignore if you only renew certs.
2027Let's Encrypt production MTCs, if the ecosystem is ready. Plan ACME client upgrades.
~2029–2035Broader industry and government deprecation windows for RSA-2048 / P-256.

The quantum transition is a change to the machinery under TLS, not a reason to stop using Let's Encrypt or to hand-issue certificates. Turn on hybrid key exchange this year. Let the CA, browsers, and ACME clients do the authentication migration on the published staging-then-production track.

This post accompanies our Let's Encrypt / post-quantum TLS notes. Follow the series at amtocbot.com.

Wednesday, August 19, 2026

Attention Is All You Need, Explained Simply

We published a plain-language walkthrough of the 2017 transformer paper — queries, keys, values, multi-head attention, and why no-recurrence mattered — with a companion animated explainer.

Tuesday, August 4, 2026

How Virtual Memory Works — LearningTechBasics

LT LearningTechBasics @amtocbot

How Virtual Memory Works

Why every program thinks it owns all of RAM.

📅 2026-08-04⏱️ ~6 min read🏷️ Systems · OS

Every process runs in its own private address space, as if it had the whole machine to itself. The OS and CPU maintain that illusion by mapping virtual addresses to physical ones, page by page.

Legend — how to read this diagram

A–EComponentsthe parts involved, labelled in the diagram
1 2 3Walkthroughnumbered steps below run in order

Translating an address

  1. Virtual address. Your program uses addresses that mean nothing to the hardware directly.
  2. Page table. The OS keeps a map from virtual pages to physical frames.
  3. TLB. A small cache of recent translations avoids walking the table every time.
  4. Page fault. If a page isn't in RAM, the OS loads it from disk and retries.

What the illusion buys you

Isolation. One process can't read another's memory — different maps.

Overcommit. Programs can use more address space than physical RAM, backed by disk.

Sharing. Read-only pages (like libraries) map into many processes once.

One-line mental model:

Give each program a fake, private map of memory — and the OS quietly translates it to the real thing.

Found this useful? Three ways to go deeper — start free.
Free · Weekly One tech idea, in your inbox

The same clear explainers, delivered weekly. No spam, unsubscribe anytime.

Subscribe free →
Guide · $39 The Open-Source AI Stack

120+ page production guide: local LLMs, fine-tuning, RAG, and domain-specific AI.

Get the guide →
Consulting Need this built properly?

AI implementation, fine-tuning and RAG done right. Free 30-minute strategy call.

Book a call →

Thursday, July 30, 2026

How Race Conditions Happen — LearningTechBasics

LT LearningTechBasics @amtocbot

How Race Conditions Happen

Two threads, one variable, and a bug that only shows up sometimes.

📅 2026-07-30⏱️ ~5 min read🏷️ Concurrency · Systems

A race condition is a bug where the result depends on the exact timing of concurrent operations. It hides in the gap between reading a value and writing it back.

Legend — how to read this diagram

A · BPartiesthe two sides of the exchange
1–nOrdereach message, numbered in sequence
1 2 3Walkthroughnumbered steps below run in order

The classic lost update

  1. Both read. Thread A and Thread B both read count = 5.
  2. Both compute. Each independently decides the new value is 6.
  3. Both write. They both store 6 — but two increments happened, so it should be 7.
  4. Non-determinism. Whether it breaks depends on scheduling, so it passes tests and fails in production.

How to prevent them

Locks. A mutex makes read-modify-write atomic, one thread at a time.

Atomics. Hardware atomic operations do increment as a single indivisible step.

Avoid shared state. Message passing and immutability sidestep the problem entirely.

One-line mental model:

When correctness depends on who wins a timing race, you don't have a program — you have a coin flip.

What Happens When You Hit "Regenerate"

You tap regenerate like it's a cheap retry. The last answer sits there, almost right, and the button looks like an eraser. It isn't....