|

Claude Code Automation: My Gates Blocked 23 Bad Posts in 10 Days

The result first: in 10 days (2026-09-28 to 2026-10-07) one publish gate ran 248 times on my Claude Code–run blog pipeline and blocked 37 of those runs. That stopped 23 different posts before they went live. Over a longer window of 49 days, the same setup logged 1,243 unattended scheduled runs across 44 jobs. None of those numbers come from a model's "done!" message. They come from append-only log files that every script writes on every run.

If you use Claude Code to automate anything that publishes, posts, or deletes, here is the claim this post defends: you need gates that read the artifact, not prompts that ask the model to be careful. Below are the real logs, the 4 failures that forced each rule, and the exact steps to copy the setup.

What I actually run (and what was measured)

I operate a multi-blog platform (my-blog.org) plus secondary channels (Tistory, dev.to, note.com) with Claude Code as the main operator. Claude Code writes and edits the Python scripts, a Windows scheduler runs them, and every script appends one JSON line per run to a ledger under data/ops/. All counts below were taken from those ledgers on 2026-10-07 by parsing them line by line.

MeasureValueSource ledger
Publish-gate runs (247 Tistory, 1 dev.to)248secondary_gate.jsonl, 09-28 to 10-07
Runs passed (exit 0)206same
Runs blocked (exit 1 = defect)37same
Runs where the checker itself broke (exit 2)5same
Distinct posts checked57same
Posts blocked at least once23same
Blocked posts later fixed and passed11same
Scheduled job runs1,243runs.jsonl, 08-20 to 10-07 (49 days with runs)
Distinct scheduled jobs44same
Runs that reported failure39same
Runs that reported success with output997same
Runs that reported success with zero output207same
Logged machine time (1,009 timed runs)43.8 hourssame, duration_s
Incident write-ups (symptom / cause / fix)169trouble_shot.md
KPI cards: 248 gate runs, 37 runs blocked, 23 posts stopped, 11 fixed and re-passed
One publish gate, 10 days · data/ops/secondary_gate.jsonl, counted 2026-10-07

Why the gate blocked posts, by defect (one run can carry two defects, so these add up to 38):

Defect codeBlocked runsWhat it means
kw-demand21The keyword-demand research file for the post was missing or older than 45 days
fig-embed6Charts were rendered but never actually inserted into the post body
rel-date4Relative dates like "this weekend" that become false if publishing slips a day
kw-grade3The declared target keyword was below the 200 monthly-search bar (Naver search volume)
fig-numbers3A number in a chart did not appear anywhere in the post text
fig-count1Fewer than 2 charts in the post

The 4 failures that created the gates

Every gate exists because something went wrong while every check said "fine". These are from the incident log, not hypotheticals.

Bar chart of block reasons: kw-demand 21, fig-embed 6, rel-date 4, kw-grade 3, fig-numbers 3, fig-count 1
Why runs were blocked · data/ops/secondary_gate.jsonl, 2026-09-28 to 2026-10-07

1. 132 posts shipped with zero images, and verification passed all of them

The Tistory pipeline had published 132 posts with 0 images in the body. Every verification run passed, because the check looked at the social preview image and exempted posts that had one. The fix was a shared gate (common_gate.py) that counts charts actually embedded in the body. That is the fig-embed and fig-count rows above.

2. A post went live, then the job crashed before writing the ledger

At 05:01:20 the daily job published a post. At 05:03 it died with ValueError: I/O operation on closed file. The cause was a late import: an image module replaced sys.stdout with a new wrapper at import time, the old wrapper was garbage-collected, and it closed the shared buffer. The ledger never got the entry, so the next morning the job would have published the same post again. The fix: switch that module to sys.stdout.reconfigure(...) and write the ledger line immediately after each publish, not at the end of the batch.

3. A --dry run silently rewrote an input file

A "dry run" of a social-follow script overwrote its target list with 30 candidates the owner had already rejected. The next scheduled run would have followed them, and dev.to has no unfollow API. The rule that came out of it: dry means no state changes, not "different output". If a dry path writes even one file, it is not dry.

4. Code pasted through a shell heredoc broke the publisher twice

When a new step was inserted into the publisher via a Python heredoc, \n inside an f-string became a real newline (SyntaxError, the publisher stopped), and \b in a regex became the 0x08 control character, so the image counter always returned 0. The second bug passed syntax checks and only showed up as "2 charts inserted, 0 images counted". Now code that contains backslashes is written through the file-edit tool, and a control-character scan runs afterward.

Why gates work where prompts don't

A prompt like "be careful, verify your work" asks the model to grade itself. Claude Code is good at writing a checker; it is not a reliable witness to its own run, because the run's log and the run's result can disagree. Case 2 above is exactly that: the post existed, and the record of it didn't.

A gate separates the two. It is a small script that opens the output (the HTML, the chart files, the ledger) and returns an exit code. The publisher refuses to continue on a defect. Claude Code then has a concrete, machine-readable reason to fix something, instead of a vague instruction to try harder. Of the 23 posts the gate stopped, 11 were fixed and re-passed. The other 12 never passed the gate in this window, so they never went out through the publisher with the defect.

Copy this setup: the exact steps

  1. Give every checker three exit codes. 0 = clean, 1 = defect found, 2 = the checker itself is broken (no targets found, file unreadable). Treat "found zero things to check" as exit 2, not as a pass.
  2. Append one JSON line per run to a ledger file: timestamp, target, exit code, defect names. If the line count doesn't grow, the check didn't run. "No new lines" is not "no problems".
  3. Make the publish step call the gate and stop on exit 1. Put the gate inside the publisher, not in a separate checklist someone has to remember.
  4. Write the ledger the moment the irreversible action happens. Not at the end of the batch. A crash between the action and the record is how duplicates happen.
  5. Tell Claude Code what a completion report may contain. In this repo's CLAUDE.md, a report can cite only three things: the checker's name and exit code, what a human actually opened, and what the gate cannot see. "Verified" without one of those is not accepted.
  6. Log every failure with symptom, cause, and fix, then turn the fix into a check. The 169 entries in trouble_shot.md are where most of the gate rules came from.

Cost and time

The scheduled jobs logged 43.8 hours of machine time across 1,009 timed runs in 49 days. That is the runtime of the scripts themselves; the gates are a small part of it. I did not log how many hours it took to write each gate, so I won't give you a build-time number. What I can say from the ledger is the trade: 23 posts stopped in 10 days is 23 posts a human did not have to catch after they were already live.

You can start with one gate today

You don't need 44 jobs to get value from this. If you have one Claude Code script that publishes, posts, or deletes, you can wrap it with a single checker that reads the output and exits 1 on the defect you fear most. Then try this: ask Claude Code to make the checker append one JSON line per run, and to stop the publish step on exit 1. Your next completion report will have something to point to besides "done".

Limits and honest caveats

Donut chart of 1,243 scheduled runs: 997 success with output, 207 success with zero output, 39 failures
1,243 unattended scheduled runs · data/ops/runs.jsonl, counted 2026-10-07
  • A gate with bad measurement lies with confidence. The first live run of a page checker flagged 320 defects across 80 pages, 4 per page. Every one was an analytics beacon failing because the checker opened pages from file://. The pages were fine; the checker was wrong. When every target fails the same way, suspect the gate first.
  • Gates see surface patterns, not quality. A title-hook detector in this repo rated a perfectly good title as having zero hooks because its word list was incomplete, and it blocked a publish until the patterns were extended. Expect false positives and keep an explicit override.
  • "Success with zero output" is ambiguous. 207 runs reported success and produced nothing. Most were legitimate (for example, a comment watcher with no new comments), which is exactly why an exit code alone can't tell you whether a job is healthy. You have to decide per job what "produced nothing" should mean.
  • This is one operator's setup. These numbers come from a single blog platform over a few weeks. They show the mechanism works here; they are not a benchmark for your project.

Found this useful?

One link is all it takes to pass this on.

Naver Blog

Comments

Comments (0)

Leave a Comment

← Back to List