In the first monthly read of this site, 274 of 440 pageviews went to a guide for fixing a Claude Cowork virtual-machine error. My consulting work is about independent reviews and implementation of AI systems. That gap suggests a question worth testing: are the people finding the site looking for help I offer? The traffic record cannot identify those readers, and the inquiry record has gaps, so I cannot call it a proven audience mismatch. It does change my next content decision: test one article aimed at a buyer’s actual implementation problem before judging the consulting offer from the current visits.
Austin Chen’s The AI Marketing Stack helped me frame that decision. It connects four jobs: learn what an audience needs, create something useful, distribute it where they look, and make a relevant offer. Then read the response and change the next attempt. Comparing that model with my site and a separate DemandForge plan led to a monthly measurement reader and its first report, plus specific content and distribution questions. It gave me no reason to buy another marketing stack.
The dated numbers below are snapshots, not a claim about today’s traffic. I paraphrase the book’s rules with chapter references and identify my own adaptations.
Find the failing layer first
The four layers help because they prevent me from fixing the step closest to a sale every time sales are slow. Chen’s chapter 10 diagnostic separates two cases for service work. If the right people engage with the content but never inquire, the offer or next step may be wrong. If those people aren’t engaging in the first place, changing the selling page cannot address the earlier audience problem.
That first snapshot covered August 12 to September 10. The Cowork guide’s traffic dominated the 30-day read, while the Work page had been seen only 13 times in the 90 days measured to August 7. The guide answers a specific software fault; it does not show whether any of its readers also need consulting help. The audience diagnosis remains a hypothesis.
A buyer-oriented article should answer a question someone would ask when an AI implementation stalls or needs outside review. I would choose the wording against actual query evidence, publish one useful piece, and watch whether people with that problem find it before rewriting the Work page again. That article has been proposed. The site has since fixed concrete problems found in buyer testing of the Work page, which was worthwhile on its own terms. It did not test whether enough relevant people reach that page.
Close the loop with a decision
Our first book analysis in August used the public catalog description because the text was not available to that review. A full-book comparison with DemandForge followed on September 8, and we applied that reading to this site on September 10. The complete reading corrected the earlier pass’s biggest criticism: the book does contain a feedback loop. Chapter 7 calls it the Signal Loop. The catalog pass still helped measure the site and reject another tool purchase, but its criticism of the book was wrong.
The book’s Signal Loop, in chapters 7 and 11, asks for a regular read of what happened and one specific adjustment. That is a more demanding contract than collecting analytics. A chart can say that visits fell. A loop has to say what evidence would separate plausible causes, what to try, and how the next read will judge the try.
We had collection before we had that decision. Google Search Console exports were archived, and GoatCounter recorded pageviews and referrers. A September review found that traffic to the Cowork guide had fallen sharply, especially from Bing, and left the cause open. The first full-book review pointed to the missing handoff: the numbers were computed, but no recurring process consumed them to select the next action.
That part was built. The monthly Signal Loop work added a reader that combines the existing traffic and search records into a dated snapshot. Its September report compares the preceding 30 days with the 30 days before them and proposes one next step. The first snapshot showed 440 site pageviews against 876 in the prior window, and Bing referrals of 107 against 350. A live Bing result-page check still found the Cowork guide first for one short error query and a longer variant on the day of the read. That single rank check says nothing conclusive about its rank during the previous month, and referral counts alone cannot prove that search demand fell. Bing’s own impression history is the missing evidence. The report therefore proposed connecting Bing Webmaster Tools so a later read could separate fewer searches from fewer clicks or a ranking change. I have no completed sign-in or second monthly verdict to report here.
For top pages and site-wide referrer families, the reader flags comparisons when either side has fewer than 20 events. The monthly report instructions prohibit recommendations from flagged rows. Chen also cautions against reading patterns from a thin body of content. This event threshold is my local reporting safeguard, not a statistical guarantee or a rule from the book.
Match the package to discovery
Chen treats the title and description as a separate decision from the subject. A search reader may type the literal words of a problem; a browsing reader may respond to a clear promise. One title doesn’t necessarily serve both routes. In a Google Search Console snapshot for July 31 to August 29, variants of the Cowork error made up 90 of 219 distinct named queries. The guide’s title closely matches what people type. Four other troubleshooting posts did not repeat its results, which argues against treating the article’s general format as the cause.
For a new informative post, I’d write down the exact query it is intended to answer and the source of that wording. If there is no known query, that is fine; it means I should judge the post against another discovery route. This is a proposed addition to my copy brief, not a measured improvement in search performance.
Chapter 5 makes a similar point about the profile people reach after seeing the work. A title that only names my job or career asks the visitor to infer whom I help. The Work page already states its offers, and buyer testing led to a clearer signal on the home page that I take client work. The book suggests checking the opening of every surface in that journey, including a social profile, for the same audience, outcome, and next step. I have not audited my current LinkedIn headline for this post, and a profile edit cannot create an audience by itself.
What changed in DemandForge
DemandForge was a separate venture exploration, not this site’s consulting offer. One historical exercise used a free guide to help a prospective garment-sewing pattern buyer compare public listings before choosing a pattern or print format; it was a sample resource, not a live offer. The full-book comparison with its venture plan went further than this site’s review. It agreed with the four layers and retained the book’s emphasis on audience intelligence, a recognizable voice, useful work, an email relationship path, and feedback. It also named where our plan departed from Chen’s examples. We proposed fewer initial assets and one primary channel, and wanted AI to finish creative work that the book often leaves at a draft for human finishing. Those were bounded design choices, not evidence that lower volume or mostly automated finishing would work better. The comparison asked for real-user calibration and a human-assisted comparison if quality or trust defects appeared. It did not validate a market, ship the workflow, or establish better economics. This comparison was a plan review, not proof of a running marketing program. An earlier DemandForge direction had been wound down before the September redesign; the comparison alone does not establish the present state of the later venture.
There was implementation after the plan review. An early worksheet and sample claimed to follow the book’s content method, but a September 8 audit found they had been composed from our task instructions without actually running a filled book prompt or doing the separate Chapter 6 voice edit. The distinction mattered: the resulting sample put repeated audit qualifications ahead of the reader’s practical decision, and an example calculation stopped before giving a usable result. We corrected one sample with an adapted Chapter 8 resource prompt and a separate editorial pass. The revised guide put a direct sleeve comparison beside the decision, separated two fabric-chart rows correctly, and completed a hypothetical print-cost calculation with assumptions visible. That is a concrete improvement in an inspected example, not evidence that novice readers found it useful or that anyone bought anything.
The book-prompt workflow in DemandForge then made that distinction explicit. It records whether a stage has a source prompt, a method, or no applicable book instruction. An applicable prompt has a saved filled input and separate venture-specific context, followed by raw output, an edit, and an independent review. Fresh contexts and retained receipts are my acceptance rules for the agents, not requirements Chen wrote for software. The same merged work saved a local job store and viewer: a worker can claim a job, preserve its inputs and outputs, submit the exact output for separate review, and render reviewed work as local HTML. The command interface itself does not generate prose or contact a customer. Its evidence is synthetic and mechanical. The repository later recorded a different current direction, so this describes the September restart work rather than an active venture plan or a live campaign.
There are real prompt runs as well as the failed initial sample. A corrected Chapter 8 resource prompt produced one sample, raw opt-in copy, and three raw welcome emails. The email copy was not edited or launched. Seven later Chapter 1 and 2 planning prompts produced an audit, three competitor reads, topic pillars, a calendar draft, and calendar validation. Those seven were sequential runs in the same agent context. They establish that filled prompts produced planning output, but do not meet the later requirement for a fresh context at each stage. The calendar was not a publishing schedule that had been executed. Later Chapter 8 distribution, scheduling and engagement prompts, a Chapter 9 paid-product prompt, and the Chapter 7 feedback prompt remained deferred in the run record. A prompt can be faithful to a source and still produce the wrong thing for a customer; those are separate checks.
Keep production and relationships separate
Two details in that comparison are easy to lose. Chapter 8’s free resource is meant to solve one small problem quickly, with a useful first result in about five minutes, followed by messages that deliver it, add insight, and invite a reply. The plan kept that purpose while allowing a direct path to the offer for someone ready to buy. The book does not demand that every buyer subscribe first. On this site, an email resource attached to the Cowork guide would mainly collect troubleshooting readers, and it would create a recurring obligation or automation cost. We have not launched one. A small, timed reply experiment remains an option if we can learn whether any of those readers have an adjacent AI implementation problem.
Chapter 5 is sharper about relationships than production. Content and scheduling can be automated; a genuine conversation with someone who replied to the work cannot be manufactured by the same pipeline. That matters to a one-person practice. Private outreach briefs and a testimonial request had been prepared when the book was reviewed, but preparation was not a sent message or a new relationship. I don’t count either as a marketing result.
My instructions to the agents
These are selected instructions I gave the agents, reproduced from the task record. They are not Chen’s purchased prompt templates. The September 5 and 6 lines came through voice transcription, including rough punctuation and the misheard name “demand forage.” The September 8 lines were typed. Their dates matter because an instruction to use a prompt does not establish that the prompt actually ran.
- September 5, reset the objective: “No, the goal was to replace whatever DemandForge was before, with uh what was described in the book, and not to keep anything from before, unless it was really valuable”
- September 6, set the redesign method: “So I want you to look at the full book copy that we just got built in the Cloud Code session, and in this session, I wanna have an independent first principles, redesign of demand forage following as closely as possible the principles in the book, unless we know for a fact something is better, that we can do better. But besides those things, we should follow the book as closely as possible”
- September 8, compare the plan: “Let’s also make sure we compare our final plan against the concepts in the AI stack book to see where we differed and make sure it’s for the best.”
- September 8, stop and audit the copy: “this is where i want to stop and make sure we are using the prompting from the ai stack book, because a lot of this text looks very AI generated and not what the book suggested.”
- September 8, require the source prompts: “yes make sure we use the exact prompts from the book at every stage”
- September 8, require actual execution: “i hope it’s actually running those prompts in fresh context sessions, not just interpreting it inline”
The September 8 copy audit in instruction 4 gives one traceable example. The initial pattern worksheet had been written from our task instructions without a filled book content prompt. For the correction, an agent ran an adapted Chapter 8 free-resource prompt. Its input named a provisional reader comparing garment-sewing PDF patterns before purchase, one small decision about a short sleeve, checked public listing and fabric-chart facts, unresolved size and print-page details, and an explicitly hypothetical printing rate. There was no measured audience pain point, email list, or established best-performing post to fill those parts as facts.
The run produced a resource draft, opt-in copy, and three welcome-email drafts. A separate agent then applied the Chapter 6 editorial method to the resource only. The edited guide answered the sleeve question directly, kept the two size-12 fabric rows distinct, and showed a complete example: 26 pages at an assumed ten cents each cost $2.60; 48 cost $4.80. The raw draft had already supplied that calculation. The editorial correction made the choice and its remaining size-group check easier to find, while correcting the raw draft’s loose fabric-row label. That is a prompt-to-output and edit record, not a timed reader test or a sale.
The later workflow requires fresh contexts and retained receipts for applicable stages; it does not retroactively make the seven historical planning runs isolated. Reading a concept, saving a prompt, running it, and seeing a customer use the result are different steps. The purchased templates, filled private inputs, and raw private outputs are not reproduced here.
The source trail for this account is the initial catalog-level analysis, the full-text site comparison, and the Signal Loop implementation.
What I would test next
The book gives me a sequence for this site:
- Read Bing’s impression history beside referrals. The first report proposed connecting the account; no historical impression read is recorded. It could show whether fewer people searched, rankings moved, or clicks changed.
- Publish one article for a buyer’s actual problem. I would choose the topic and wording from query evidence, then ask whether relevant readers reach the consulting work. No article has been published from this recommendation.
- Check the route from that article through the profile to the offer. If relevant readers arrive, I can look for where they stop. Current buyer traffic is too thin for a conversion verdict.
- Consider a bounded email and reply test. If the earlier tests identify a relevant segment, a useful resource might show whether it starts named conversations. This remains optional.
The useful leading indicators are qualified engagement and inquiries, with enough evidence to know which audience generated them. The book’s creator revenue tables are not forecasts for this consulting work. If a test brings no relevant visits, I should revisit audience and distribution before polishing the offer again. The monthly reader and its first report give me a way to choose the next test; a later read still has to judge what happened.