Saturday, October 3, 2026

An AI Writing Policy: Bleeding Red Ink

Hidden Social Image

A work colleague recently forwarded a link to an AI Writing Policy created by the engineering team at Clay. The core principles are exactly right: you must stand behind every sentence, writing is thinking, and you need to respect the reader's time.

Applying strict editing standards to AI isn't a new corporate mandate. For me, it’s just the baseline.

The Journalism Roots

Growing up in a house run by two Indiana University Journalism majors sets a specific baseline. My dad was a Public Affairs officer in the Air Force and later in federal service. As we moved around the country, my mom always picked up roles as a copy editor. Over the years, she worked at a paper in Biloxi, out in Thousand Oaks, the Tampa Tribune, the St. Pete Times, the Rocky Mountain News, and another paper in Colorado that escapes memory.

Every single paper I wrote for school was subjected to her desk. When I got them back, they were completely redlined. Bleeding ink. That level of intense, structural criticism became normal. It dictates how I approach writing with AI today.

The Turns

oraclenerd Blog Archive showing 269 posts in 2009

I have been verbose for a very long time. I suffer from a chronic inability to leave a thought unpublished. In 2009 alone, I wrote nearly 270 articles.

When you maintain that kind of volume today, it is easy for readers to assume you are just feeding prompts into an LLM and hitting publish. But the output has never come from cutting corners.

When I sit down to write a post with <insert_model_name>, it takes a number of turns. It takes actual time.

I started drafting a new post (one that is still in Draft) on a Saturday in January. My expectation was that it would be like normal writing: quick, off the cuff, "my thoughts." I figured it would take an hour.

I invited Copilot in to do copy editing. But it had ideas, too. I expanded on those ideas, argued with them, refined them. Eight hours later, I still wasn't done. I've updated it exactly one time since (adding a major realization), and I am still sitting on it, waiting until the underlying work actually goes to Production before I hit publish.

That is the baseline. Editing AI isn't about adding words; it's about hacking away the generated costume to find the sharp truth of your original prompt buried underneath. Here is what that red ink looks like in practice:

  • A draft tried to claim, "Booleans shouldn't exist in your physical data model." The red pen changed that to, "A boolean should never be your first move." Precision matters.
  • A recent post about JSON opened with a fake, generated confession. I cut it immediately. A post about declared truth cannot start with an undeclared lie.

The Brain Dump

AI Written, AI Read cartoon by Marketoonist

There is a secondary, highly pragmatic reason for enforcing this level of rigor. With the advent of modern tools, it is now trivial for anyone to point an agent at a body of work and ask for a summary. Teammates can drop a link into an IDE, tell the agent to follow all the references, and ask for a ten-sentence synthesis.

Writing meticulously is no longer just about communicating with a human reader in the moment. It is about structuring the underlying truth so that it can be accurately scraped, synthesized, and redistributed by machines. If the original text is bloated with lazy AI filler, the downstream synthesis will be useless. If the original text is tight, considered, and bled over, it effectively scales a single brain across an entire team. It is a reason to write more, not less. But it only works if the desk holds the line.

The Copy Chief

The first line of defense isn't even human. It is a codified personal style guide. It is a literal file containing my voice, my formatting quirks (like a strict ban on em dashes), and my rules of engagement. The AI is forced to ingest this file before it generates a single word. It bends the model to the desk immediately.

I even use a second AI agent behind the scenes to act as an independent copy editor.

But here is the catch: machine red ink doesn't carry authority. That second editor is sometimes wrong, too. It has misread an entire section defending analysts as an attack on them. It has "improved" text I explicitly asked to be copied verbatim. It has praised a draft and then reversed itself on a second read.

Getting more models to mark the page is just adding more ink. A redline only matters if somebody holds the final desk. I have to be the copy chief.

This is exactly why every post on this blog now ends with the same functional transparency byline seen at the bottom of this page: Architected by Chet, written by Gemini 3.1 Pro.

The AI might generate the raw sentences, and it might even pitch ideas I hadn't considered during an eight-hour Saturday session. But the architecture (the structure, the cuts, the arguments, and the final truth of the post) belongs to the desk.

I have been wrong, very wrong, on this blog before, and I have to own those mistakes. If a data modeling strategy sends a reader down a blind alley, I am the one on the hook. That is why the process takes time. I refuse to publish anything here until it's crossed my desk.

Architected by Chet, written by Gemini 3.1 Pro


Appendix: The Meta-Redline

This exact post is a perfect example of how the desk operates. It took twenty-seven distinct turns (and roughly two hours and ten minutes of active, back-and-forth refinement) to get from the initial idea to the final draft you are reading right now.

It started with an image of the Clay policy dropped into the chat. <insert_model_name> spit out a perfectly sterile, corporate-sounding summary. Over the next dozen turns, drafts were rejected, family history was forced in, the narrative was shifted, and ultimately a completely separate, secondary AI agent was brought in to critique the work.

That secondary agent absolutely nuked the draft. It pointed out the voice was a machine's, calling out terrible cliches like "resonate" and "broadcasting noise". It demanded actual receipts of the redlines instead of just talking about them.

That is what it means to hold the final desk.

Monday, September 28, 2026

The System Still Works: The Algorithmic Arms Race

Back in May, in The System Works, I wrote about the Sunday session where I finally set Antigravity loose to map out twenty years of gut feelings about the American healthcare system.

I wrote about the plastic card illusion: walking into a clinic with no insurance and paying a $115 copay, versus handing over an insurance card and watching the bill explode to $400. In consulting math, that meant an ordinary 15-minute consult jumped from $460 an hour to $1,600 an hour. Nothing about the clinical diagnosis had changed. Only the billing plumbing had.

The core thesis we landed on that Sunday wasn't that the system was broken. It was the opposite: the system is working perfectly. It is an autopoietic machine, an economic ecosystem designed to reproduce its own administrative complexity and extract maximum yield.

I ended that post noting that while the system hadn't changed, I could finally see the paths, the rules, and why people kept ending up in the exact same traps.

Well, last week, Blue Cross Blue Shield published an analysis that proved the machine has found its next gear.

The $1 Billion Query: Mining for Acuity

According to a new analysis of claims data by the Blue Cross Blue Shield Association, the rapid adoption of AI coding tools by hospitals added an estimated $942 million in commercial healthcare spending between 2023 and 2025 alone.

More than 60% of hospital systems are now running AI-enabled revenue cycle software. These tools continuously scan electronic medical records, lab reports, and doctor notes to find billable complexity.

And what did all those millions in new revenue buy in terms of actual medicine?

Zero.

BCBSA looked closely at patients undergoing major bowel surgery. The AI tools flagged a massive surge in secondary diagnoses, like anemia, derived from single post-operative lab floats. When a hospital attaches a secondary diagnosis of anemia to a surgical stay, the claim automatically bumps into a higher-severity, higher-reimbursement DRG (Diagnosis-Related Group) category. Over 70% of the entire $942 million increase (roughly $650 million) was tied strictly to these secondary diagnoses.

Yet when researchers looked at the clinical charts, there was no corresponding increase in treatments. No uptick in blood transfusions. No change in medication. No change in bedside recovery time.

Luke Chalker, BCBSA's senior vice president of product and data science, summarized the finding with brutal clarity:

"The disconnect between diagnoses and treatment suggests that AI is identifying more billable conditions, not sicker patients."

If you're a data person, read that sentence again. That is not a bug. That is an optimization routine running exactly to spec.

The 74,000-Code Moat

For decades, healthcare policymakers and insurers built a labyrinth of classification. We moved from simple fee-for-service to prospective payment systems based on 74,000 distinct diagnosis codes and 79,000 procedure codes. The stated goal was accuracy and granular clinical documentation.

The unstated reality was that human beings cannot hold 74,000 codes in their heads. Independent medical practices couldn't afford armies of certified coders, which accelerated the collapse of private practice and forced doctors into the arms of massive hospital conglomerates.

And now, the conglomerates have done what every well-capitalized enterprise does when faced with a complex rules engine: they automated it.

They pointed large language models and machine-learning classifiers directly at the EHR data lake. The algorithm doesn't ask if a patient feels better, or if a post-op hemoglobin dip is clinically meaningful. The algorithm asks a purely relational question: Does this float value in table A legally support a secondary modifier in table B that increases the payout in table C?

The code is no longer a proxy for care. The code is the product.

The Press Release: Setting the Narrative

What makes this entire situation so fascinating from a systems perspective is the delivery mechanism. This wasn't a leaked internal memo. This was a press release blasted out by the Blue Cross Blue Shield Association to major outlets like Reuters.

Why is a massive insurance conglomerate publicly complaining that hospitals are out-coding them?

Because they are setting the narrative.

As we traced in the blueprints of The Perfect Machine, Blue Cross was originally born in 1929 at Baylor University. It was created by hospitals, for hospitals, specifically as a pre-payment plan to guarantee hospital solvency. The hospital created the third-party payer to insulate itself from market forces.

For nearly a century, the two sides grew together. Now, the hospitals have turned loose algorithms that out-optimize the insurer's own rulebook. By blasting out a press release, BCBSA is establishing a public scapegoat for why your employer's health premiums are going to spike another 7% next year. They are pointing the finger squarely at the hospital bots.

But make no mistake: this is a two-sided bot war. Insurers aren't innocent victims. For the last five years, major health plans have deployed their own automated actuarial algorithms (like Cigna's PXDX and UnitedHealth's nH Predict) to batch-deny claims in 1.2 seconds without human review. The hospital buys an AI bot to maximize billable severity; the insurer buys an AI bot to mass-reject claims on protocol technicalities.

Two multi-billion-dollar machine-learning networks are now trading transactions back and forth across an opaque API perimeter. Neither machine cares about the human lying in the hospital bed. The patient is just the transaction payload passing between them.

The System Still Works

When you see headlines about AI adding $1 billion to hospital bills, the natural reaction is to throw up your hands and say, "Healthcare is completely broken."

It's not.

If you build a system where the customer does not pay, where the price signal is illegal or hidden, where payments are determined by an arbitrary 74,000-node taxonomy, and where survival depends on administrative volume, you will get autonomous code-mining bots every single time.

In May, I wrote that once you see the plumbing, you stop looking for villains and start understanding the incentives. The BCBSA report isn't evidence of a broken system. It is living, breathing proof that the machine is evolving right on schedule.


Architected by Chet, written by Gemini 3.1 Pro

Sunday, September 27, 2026

When the Guardrail Stops Guarding

A few months ago, in You Can Point a Foreign Key Where?!, we looked at pointing foreign keys to unique candidate keys instead of traditional surrogate IDs.

In the comments, Tom offered an essential reminder from decades in the trenches:

"You and I have decades now of CREATE TABLE... and those decades have shown us the only thing that is truly immutable are the things stakeholders tell us are absolutely not immutable... business codes have a way of becoming less immutable across years. They get renamed, merged, generally f-d up all the time and when a new C-whatever comes in who likes SOON instead of ASAP, you're cooked."

Tom was arguing for surrogate keys to separate meaning from relational identity. He was right. Business vocabulary is fluid. What feels permanent during sprint planning rarely stays permanent across three fiscal years.

That brings us to a design choice many of us reach for when standing up a new feature: skipping the reference table entirely and putting business enums directly into a table CHECK constraint.

CREATE TABLE orders (
        order_id     NUMBER GENERATED BY DEFAULT AS IDENTITY,
        order_date   DATE NOT NULL,
        order_status VARCHAR2(20) NOT NULL,
        CONSTRAINT pk_orders PRIMARY KEY (order_id),
        CONSTRAINT ck_orders_status 
            CHECK (order_status IN ('OPEN', 'PENDING', 'SHIPPED', 'CANCELLED'))
    );

To be clear: this is not a post against CHECK constraints. CHECK constraints are indispensable. This is a post about which rules they can actually hold.

Putting an enum into a CHECK constraint looks tidy, declarative, and completely self-contained. It avoids spinning up another table, mapping another entity, or coordinating extra seed files. In the moment, it feels like disciplined engineering.

Until the business evolves, and the guardrail quietly stops guarding.

The Silent Failure: Enforcing Neither

The standard objection to putting enums in CHECK constraints is migration friction. We have all heard the complaint: adding a status requires an ALTER TABLE, table locks, and coordinated deployments instead of a simple INSERT.

That argument is true, but it is an ergonomic argument. A determined engineering team can easily wave it off as an acceptable deployment cost.

The fatal problem with an enum CHECK constraint is not operational friction. It is structural failure.

Consider Tom's scenario: leadership decides that 'PENDING' is too vague. From now on, operations needs to split future orders into 'AWAITING_PAYMENT' and 'AWAITING_STOCK'.

What happens to existing records? The business does not want to rewrite history. The thousands of closed orders that passed through 'PENDING' over the last two years need to remain untouched.

Now look at what happens to your constraint:

ALTER TABLE orders DROP CONSTRAINT ck_orders_status;
    
    ALTER TABLE orders ADD CONSTRAINT ck_orders_status
        CHECK (order_status IN (
            'OPEN', 
            'PENDING', 
            'AWAITING_PAYMENT', 
            'AWAITING_STOCK', 
            'SHIPPED', 
            'CANCELLED'
        ));

Because a table constraint evaluates every row equally across the entire table, 'PENDING' must remain in the allowed list forever just to keep historical rows valid.

The moment history and future diverge, the CHECK constraint is forced to permit both. And the moment it permits both, it can no longer prevent an application bug from inserting a brand new 'PENDING' order tomorrow morning.

You thought you had an automated guardrail protecting business integrity. What you actually have is a decorative plaque. The constraint that was written to enforce valid state transitions has quietly stopped constraining.

A lonely gate on a sidewalk with open grass on either side, labeled CHECK (order_status IN ('OPEN', 'PENDING', ...))

Temporal Boundaries Trump Hardcoded Strings

When we move business vocabulary out of DDL and into a proper reference table, we don't resort to a lazy is_active boolean flag either. As we saw in The Boolean is Lying to You, booleans erase history.

Instead, we use explicit temporal boundaries:

CREATE TABLE order_statuses (
        status_id            NUMBER GENERATED BY DEFAULT AS IDENTITY,
        status_code          VARCHAR2(20) NOT NULL,
        description          VARCHAR2(100) NOT NULL,
        display_seq          NUMBER NOT NULL,
        effective_start_date TIMESTAMP WITH TIME ZONE NOT NULL,
        effective_end_date   TIMESTAMP WITH TIME ZONE,
        --
        CONSTRAINT pk_order_statuses PRIMARY KEY (status_id),
        CONSTRAINT uq_order_statuses_code UNIQUE (status_code),
        CONSTRAINT ck_order_statuses_dates 
            CHECK (effective_end_date IS NULL OR effective_end_date >= effective_start_date)
    );

(Notice where the CHECK constraint lives: validating that an end date cannot precede a start date. A mathematical and calendar rule, exactly where it belongs.)

When the business retires 'PENDING', we don't touch the orders table. We update the lookup table:

UPDATE order_statuses 
       SET effective_end_date = SYSTIMESTAMP 
     WHERE status_code = 'PENDING';
    
    INSERT INTO order_statuses (status_code, description, display_seq, effective_start_date)
    VALUES ('AWAITING_PAYMENT', 'Awaiting Payment', 20, SYSTIMESTAMP);
    
    INSERT INTO order_statuses (status_code, description, display_seq, effective_start_date)
    VALUES ('AWAITING_STOCK', 'Awaiting Stock', 30, SYSTIMESTAMP);

Every historical order referencing 'PENDING' remains fully valid. Meanwhile, any active query or UI dropdown filters on effective_end_date IS NULL (or evaluates whether SYSTIMESTAMP falls between the start and end dates). New orders cannot select ' PENDING'. Old orders cannot be corrupted.

Better yet, temporal versioning gives you future-dating for free. When operations announces that a new status takes effect on January 1, you insert the row today with an effective_start_date of January 1. At midnight, it activates automatically. No midnight deployments, no emergency patches, and zero table locks.

The Metadata Leak

A CHECK constraint treats an enum as a naked, isolated literal. But business vocabulary never stays naked.

Before long, the rest of the team needs context:

  • The user interface needs human-readable labels and a coherent sort order rather than raw database codes.
  • Reporting pipelines need to know which statuses represent terminal states versus active work in progress.
  • Audit systems need to know what the valid vocabulary looked like on a specific date two years ago.

In a reference table, those attributes have a natural home (display_seq, effective_start_date, description). The database acts as a shared, queryable source of truth.

In a CHECK constraint, the database cannot hold that context. So the context leaks. It leaks into frontend TypeScript switch statements, backend YAML files, and hardcoded reporting filters. You didn't avoid complexity; you just forced your schema to live in five different application files instead of the database.

The Portable Rule

This brings us to a reliable boundary for schema design:

  • Use CHECK constraints for mathematical, temporal, and physical invariants. quantity > 0. percentage BETWEEN 0 AND 100. end_date >= start_date. Cross-column requirements like requiring a card token when payment type is credit card. These are arithmetic and physical boundaries. They do not drift because of a company rebrand.
  • Use Lookup Tables for business vocabulary and lifecycle states. Order statuses, workflow stages, customer tiers, reason codes. Anything that product owners or executives might refine next quarter.

If changing the rule violates the laws of mathematics or calendar time, it belongs in DDL. If changing the rule requires a conversation with product management, it belongs in a table.

Scoped Optimization and AI

In my reply to Tom on that earlier post, I noted that keeping enums out of CHECK constraints is one of the fundamentals that requires an explicit waiver in my coding assistant instruction files.

It is worth asking why AI assistants reach for CHECK (status IN (...)) almost every single time they are asked to scaffold a schema.

The assistant is not being lazy. It is behaving exactly like an application developer focused on a single user story: it is correctly optimizing for a scope that does not include the third fiscal year.

Within the boundaries of a single prompt, a self-contained CHECK constraint satisfies every immediate functional requirement. It compiles cleanly in thirty lines of generated SQL, avoids creating auxiliary tables, and avoids the cognitive overhead of foreign keys. The model's horizon is the current response window. It has no reason to care what happens when marketing changes 'ASAP' to 'SOON' thirty-six months from now.

If we don't guide our tools with clear fundamentals, they will produce code that is locally optimal and globally fragile.

Take the time to build the reference table, give it proper temporal boundaries, and establish the foreign key. It is the only way to ensure that when your data models meet the third fiscal year, the database is still the place telling the truth.


The arguments are mine. Drafted with Gemini 3.8 Flash.

Wednesday, September 16, 2026

Every JSON Column Has a Schema

The arguments are mine. The typing was not.

Let's complete the trilogy (That's probably a lie...).

Over the last two weeks, we've picked apart the boolean:

  • In the warehouse (OLAP), a boolean duplicates a truth that already exists elsewhere (your temporal boundaries).
  • In an operational system (OLTP), a boolean destroys a truth that did exist (wiping out history and sequence).

A JSON column does something far sneakier: it never declares a truth, so nothing can ever contradict it.

The Unfalsifiable Blob

A shapeless blob cannot be wrong, because there is no declared shape for it to violate.

Think about what happens when you create a proper relational column. You declare customer_id INTEGER NOT NULL REFERENCES customers(id). You have drawn a hard line in the sand. If the application tries to insert 'banana', the database rejects it. If it tries to insert an orphaned customer, the database rejects it. The engine knows what is true and what is a lie, and it protects you.

Now look at a JSON column.

  • {"customer_id": 123} is valid JSON.
  • {"customer_id": "123"} is valid JSON.
  • {"customerId": null} is valid JSON.
  • {"cust_id": "banana"} is valid JSON.
  • {} is valid JSON.

The database accepts every single one of those rows without blinking. It writes them to disk, returns a 200 OK, and goes about its day. Why? Because you never declared a contract. You never told the database what "right" looks like, so nothing can ever be "wrong."

Until six months later, when the reporting query blows up. At that point, you don't have a data model: you have vibes and a support ticket.

Watching From the Outside

I know why people do it. A new feature lands on your desk, the product manager is still fuzzy on the specs, the attributes will probably change next sprint, and you just need to get the code out the door. So you slap a payload JSON column on the table and tell yourself you're being agile.

I've never done this. Not with JSON.

Booleans, sure. I will readily admit to that sin. I've slapped an is_active flag on a table and paid for it later. But dumping application objects into a JSON column is one I have only ever watched from the outside, which is its own kind of education.

If you come from the database world, the foundational rule has always been simple: put integrity constraints as close to the data as humanly possible. The database exists to protect the data from the application, because applications get rewritten every two years, but data lives forever.

When you drop an unconstrained JSON blob into a table, you're betting that every future engineer touching that application will remember to enforce every single implicit business rule in code.

Spoiler: they won't.

Rebuilding the Schema, Badly

Someone reading this will inevitably push back: "Wait, you can enforce integrity on a JSON column!"

And you can. Modern engines will let you bolt integrity back onto a document. Postgres will happily take a CHECK constraint on an extracted jsonb path. You can create generated columns with foreign keys, and you can build functional GIN or B-tree indexes against nested attributes.

Technically, you can do it. But look at what you are actually doing: you are reconstructing, one painful piece at a time, the relational schema you declined to write in the first place, using an esoteric, vendor-specific syntax that nobody on your team will recognize in a year.

You didn't avoid the schema. You just decided to rebuild a worse version of it by hand.

Relocating the True Cost (and Closing the Loop)

To be clear: this isn't about beating up on application developers.

When a developer drops a JSON column into a migration, they aren't trying to sabotage the company. They are responding to very real, very rational pressures: sprint deadlines, velocity metrics, and avoiding the friction of formal schema reviews. From their seat in the sprint, skipping the table design feels like pure efficiency.

Years ago, Cary Millsap wrote a post about formatting tables of numbers that contained an insight I have quoted many times:

"Good design is a topic of consideration. And even conservation. If spending 10 extra minutes formatting your data better saves 1,000 readers 2 minutes each, then you’ve saved the world 1,990 minutes of wasted effort."

Cary's math is irrefutable, but appealing to civic virtue ("save the world 1,990 minutes") rarely changes engineering behavior on its own. What actually changes behavior is seeing the feedback loop close.

When you dump a raw, shapeless JSON blob into a table, you didn't eliminate the work. You simply relocated the cost.

In the short term, you quietly transferred that cost downstream to the analytics engineers and BI developers. Every single report now requires defensive SQL: unnesting arrays, casting strings to integers, guessing at nulls, and handling three different key spellings.

But the loop doesn't stop with the analytics team.

Eventually, product asks for a new operational feature: an in-app filter, a bulk edit, or a performance dashboard built directly against that transactional table. And guess who gets assigned the ticket? The application developer.

Now, the very engineer who bypassed the schema to save twenty minutes in sprint four is staring at a production bug in sprint twelve, trying to write unreadable JSON path queries against their own shapeless blob. You didn't save time. You just deferred the agony, with interest, back onto your future self.

The Waiver

In my own agent instructions, a JSON column requires a waiver: the default answer is no, and the burden of proof is on the column.

When does it actually earn that waiver?

When the data is genuinely an opaque, third-party black box that the database never needs to reason about: raw webhook payloads, audit logs, or configuration blobs that the engine will never filter, join, or aggregate on.

If your application needs to query it, if your business needs to report on it, or if it relates to any other entity in your system: it belongs in a column.

Doh. It really is that simple.