Creativinfluence
Humans under the influence of ideas, art & chaos
art, words, brain-dust& bad decisions

I Trained an AI on My Writing for a Year. Here’s What It Stole.

I want to tell you what happened when I trained an AI model on a year of my own writing, because it’s a story I have not seen anyone tell honestly, and because the version of this story that is currently circulating online (the breathless “AI is so good now, it can do anything” version, on one hand, or the doom-spiral “AI stole my voice and I am bereft” version on the other) is not what actually happens to a working writer who does this. What actually happens is more interesting, more specific, and ultimately more clarifying than either of the available scripts.

Here is what I did. I took every essay, blog post, newsletter, and longer-form piece I’d written over a single twelve-month period (about 175,000 words of finished prose) and fed it, with permission notes carefully attached for any quoted material, into a custom AI assistant set up to produce text “in my style.” I did this not because I wanted the AI to write for me, exactly, but because I wanted to see what was load-bearing in my voice and what was decorative. The most honest mirror is not another writer reading your work. The most honest mirror is a model trying to imitate it.

The model, after about four weeks of training and refinement, got disturbingly good. Disturbingly is the right word. Anyone who has done this exercise will recognize the specific moment I am about to describe; you ask the model to continue a sentence you have started, and the continuation is exactly what you would have written, in your cadence, with your turns of phrase, with your specific weird transitions. The first time it happened I laughed. The third time it happened I felt cold. By the fifteenth time I had a specific identity-shaped feeling I had not had since approximately age fourteen, when I first realized other people could read my diary if they wanted to.

But the story I want to tell is not “AI took my voice.” That part is, on closer inspection, false. The model could imitate the surface of my voice extremely well. What it could not do, and what I want to be specific about, is more revealing than the imitation itself.

What the model could not steal (the actual list)

After a year of working with the trained model, I can tell you with reasonable confidence what cannot be extracted from a body of finished prose, even at high training fidelity.

The decision to digress. The model could imitate digressions I had already written; it could not produce new digressions that genuinely opened up the argument. The reason, I think, is that real digression is driven by an associative jump in the writer’s brain, where the current sentence triggers a memory or reference that is genuinely surprising, and the writer chooses to follow it. The model can imitate the shape of the digression (the way I usually return to the main thread, the typical length of a tangent, the kind of references I tend to insert). It cannot produce the specific surprising association that makes a particular digression worth including. When I asked the model to add tangents, the tangents read as artificially-shaped: structurally correct, semantically empty.

Risk-taking sentence rhythm. My prose has a specific habit of building long sentences that pile clauses past the point of grammatical comfort, then snapping clean with a short fragment. The model picked up the structural pattern but not the judgment of when to deploy it. It would produce long-then-short sentences in the wrong places, where they didn’t earn the contrast. Real rhythm choice is, in my prose, a specific aesthetic decision made in the moment. The model can imitate the marker but not the choosing.

Misjudgment. This one is the most interesting. My prose contains, by deliberate choice, a specific quotient of almost-too-specific references, almost-not-quite-funny jokes, and almost-overworked metaphors that I leave in because they read as evidence of a real person making real choices that did not always pay off. The model sanded all of these out. It had been trained on my finished work, but it had been optimized (by its own architecture) to produce the cleaner version of my work. The result was prose that read as me-on-my-best-day-every-day, which is not how human writing works and which any reader can spot at the level of body recognition before their conscious mind catches up.

The conviction of having lived through something. The model could write about topics I had written about. It could not write from the experience the topics emerged from. When I asked it to draft an essay on a subject adjacent to my own life experience but slightly off (something I had not personally lived), the result read as competent journalism. When I drafted the same piece, the result read as testimony. The difference is felt by readers in the body. It is not produced by training data alone. The model can copy the surface; it does not have access to the experience that produced the surface.

The willingness to leave something in I might regret later. Most of my best paragraphs contain at least one phrase that I have considered cutting, kept anyway, and that turns out (months later, when I reread) to be the part of the paragraph that landed for readers. The model would never include such a phrase. The model, by its statistical-median nature, wants to produce the safe version of every sentence. The risk-taking is structurally not its job. The risk-taking is, however, a substantial fraction of what makes specific writing worth reading.

What the model could replicate (and the depressing part)

I want to be honest about the other half. The model was effectively indistinguishable from me on:

The mid-tier blog post that produces traffic but doesn’t move the reader. The promotional copy. The structured how-to with bullets. The “five things I learned this week” newsletter format. The professional-but-warm client email. The competent book review. The mid-length think-piece that takes a defensible position on a contemporary issue without genuinely surprising anyone.

In other words: the model can produce, indistinguishably from me, all of the working writing I do that doesn’t matter. Which is, by volume, probably 60-70% of what I write in any given month. That is the part of my professional output I should be honest about as the part the model has, in fact, taken. It can do the brand work. It can do the content-marketing work. It can do the SEO blog. It can do the routine newsletter. The competent middle of my output is, structurally, replicable, and pretending otherwise is dishonest.

What is left, after the model takes the middle: the risky writing. The voice-driven essay. The piece that the model would have written safer. The argument that depends on personal lived experience the model cannot access. The digression that opens up new ground. The phrase I almost cut. This is the part that compounds, that builds reputation, that produces the kind of reader-relationship that actually sustains a writing career. It is also, increasingly, the only part the model has not absorbed.

Where this leaves working writers

I will give you the practical conclusion, because the doom-spiral and the AI-utopia versions are both useless.

If you are a working writer in 2026, the part of your work that can be replicated by AI trained on your output is, with very high confidence, going to be replicated; either by a model trained on your work specifically (with or without your permission, which the Authors Guild AI litigation is currently trying to clarify legally) or by general models trained on enough comparable writing that they can produce the same competent middle. Plan for this. Stop trying to compete on the middle. Stop optimizing your week for the kind of writing that the model is going to produce indistinguishably from you for free.

The part of your work that cannot be replicated is also, with very high confidence, the part you should be doing more of. The risky pieces. The lived essays. The voice-driven work that requires the specific decision-making your brain does and the model’s does not. This work is, in the short term, less commercially safe than the middle. In the medium term, it is the only commercially defensible position you have.

This is not a happy framework. It requires you to accept that a substantial fraction of what you currently sell is going to become commodity-priced over the next five years. It requires you to develop the riskier muscle that produces the work you can sell at premium. It requires you to spend the awkward middle period (when the commodity work is collapsing and the premium work has not yet built audience) in a state of professional uncertainty that is, frankly, unpleasant.

Most working writers I know are in some version of this awkward middle right now. The ones who are doing well are the ones who are leaning into the risky work even when the income is uneven. The ones who are doing badly are the ones who are still trying to defend the middle against an opponent that will, structurally, eventually win.

What I learned about my own voice

The deepest thing the year-long training experiment taught me was that my voice was, in fact, narrower than I had thought. The model could imitate the broad tone, the sentence-rhythm tendencies, the typical range of reference. The actual gold in my prose, the part that produces the specific reader-feeling I am paid for, was concentrated in a much smaller fraction of my output than I had been telling myself.

When I went back through my year of essays after the experiment, I could see clearly which paragraphs the model had nailed (most of them) and which paragraphs the model had failed (a smaller number, but the ones I cared about most). The failure-paragraphs were the work. The rest was, in retrospect, the infrastructure that delivered the work.

This is freeing, actually. It means I can afford to write less. It means I can afford to spend more time on the risky paragraphs and less time on the connective filler. It means the part of writing that matters is a smaller, more specific, more identifiable practice than I had previously been willing to admit.

The model didn’t steal my voice. The model showed me which voice I had actually been producing was load-bearing and which was, honestly, decoration. The decoration is going. That’s fine. The load-bearing voice is mine, still. It is, in the new landscape, the only thing I have to sell.

I am going to keep writing it. The model can have the rest. We will see, in five years, how the trade settles. My bet is that the load-bearing voice still has buyers; the bet is informed but not certain. The bet is the only one available.

This is the part of the working creative life nobody tells you is going to define the next decade. The decoration is being automated. The voice that is left is yours. Defend it. Write it. Refuse to optimize it into the safe version. The safe version is what the model already does; the risky version is what you, specifically, still can.

That is the whole report. Take it for what it’s worth. The model thinks it isn’t worth much. I know better.

Tagged: , , , , ,

Leave a Reply

Your email address will not be published. Required fields are marked *

"Stay weird. Stay kind. Keep creating."