awp / autonomous agentic website pipeline

Buyer questions

Answered straight, including the ones that lose the sale.

The objections a skeptical technical buyer actually raises. Every factual claim below is reproducible by running a command in the repository. Where the honest answer is "no" or "unknown", it says so.

199 USD once. No subscription, no account, no hosted service.

Figures measured 2026-07-30 at 21:19 local. The reproducing commands are at the end of this page.

Section 1

About outcomes

Will this get my site ranked?

No. Nothing here will ever claim otherwise.

The pipeline produces zero off-site signal. The repository's own commercial analysis states its most likely failure mode explicitly: the pipeline delivers exactly what it promises, a technically correct and fully indexable site in a fraction of the usual time, and then day 90 arrives with almost no clicks, because ranking is limited by domain authority and external signal and this produces neither.

The structural version, from that analysis: this is a supply-side tool sold into a demand-side problem. Making websites was already nearly free. Getting them seen is the expensive part, and nothing in layers 0 through 12 touches it.

Will my pages get indexed?

Unknown, and not controllable by anyone.

Google indexes a subset of what it crawls, and it gives nobody a mechanism to compel inclusion. Submitting through Search Console or IndexNow is a request, not an instruction. A technically flawless site can sit at partial or zero indexation indefinitely.

Layer 11 reports measured Search Console coverage with no probability language. Zero indexed pages at day 30 is treated as a real, reportable outcome that routes into the monitoring layer rather than being smoothed over.

What results have other buyers seen?

There are no other buyers, so there are no reviews, no case studies, no revenue figures, and no measured outcome distribution. Nobody has run this enough times to produce one.

If a page for this product ever shows you a customer count or a testimonial, treat it as fabricated, because as of this writing there is nothing to draw one from.

How long until I see traffic?

I don't know, no measured base rate exists, and any number I gave you would be invented.

What is documented is the shape of the wait: elapsed time to first confirmed indexation is days to weeks and can be longer on a new domain, and the monitoring layer's first genuinely useful cycle is the one after enough impressions exist to clear its volume floors. The stop-loss rule evaluates at day 180.

Is this going to get me penalized for AI-generated content?

The risk register treats this as the highest-impact risk in the product and states the real danger precisely. It is not that one site gets actioned. It is that every site produced by one pipeline shares a fingerprint (same template lineage, same section structure, same internal linking pattern, same hosting footprint), so a pattern-level action hits every site in the same week.

The specified mitigations are design decisions, not disclaimers: an information-gain gate that fails any page restating the top five results without adding a distinct fact, dataset, tool, or perspective; a requirement for at least one non-reproducible element per site; a hard page cap per site; and never sharing hosting, analytics identifiers, or contact details across sites.

The honest caveat: no mitigation makes scaled content safe, which is why the documentation's advice when a site has no genuine differentiator is to stop before deploying rather than to generate harder.

Section 2

About the state of the product

What is actually finished?

Run ./verify.sh. It performs 14 checks and currently exits 0 with "State: consistent and complete":

  • 18 Python modules compile.
  • 13 layer contracts present, 0 missing files.
  • 204 layer gates (182 blocking), 0 parse failures.
  • Every assertion name in use is implemented.
  • Contract validation reports 0 errors and 0 warnings.
  • 4 vendored design prompts.
  • The domain engine generates, filters, and scores.
  • 7 GEO sprint stages, 0 missing files, 191 sprint gates.
  • The fleet keyword registry loads and reports 0 cross-domain conflicts.
  • The sales page contains 0 external network references.
  • A NASA Power of 10 audit reports no build-failing violations.
  • The regression suite runs 348 tests.
  • The gate and schema cross-check finds no gate contradicting its own schema.
  • Every declared external input is present.

What is not finished?

No end-to-end run has ever been completed. Every layer contract is written and every gate is evaluable, but no site has been taken from intake through indexation by this pipeline. That's the honest gap, and it's on the sales page.

Consequences you should price in: there is no demo site, there is no measured median for anything (cost per page, attempts per layer, hours at the design step, indexation rate), and the first person to run it end to end will find friction that only a real run surfaces. That person may be you.

Is there a demo site built with this?

No. Showing you a site built some other way and implying the pipeline produced it is the exact fabrication the gate suite exists to prevent.

Why does your own demo run fail its gates?

Because it should. The awp domains convenience command deliberately does not run the trademark and prior-use screens, so gate G-L03-04 fails with "20 of 20 flags are absent; cannot prove none are true", and G-L03-12 fails because the provenance fields naming the screening sources are empty.

That is the intended behaviour and it is documented in the repository as a known design tension: the shortlist command can never satisfy the evidence coverage floor on its own, by construction, because the shortlist is explicitly not cleared for purchase until the screens run.

A pipeline that passed there would be telling you a domain is safe to buy on evidence it never gathered. That is the failure this whole product exists to prevent, and it would be embarrassing to commit it in the demo.

Section 3

About value

I could rebuild this in a weekend. Why pay?

You could rebuild the visible surface quickly, and the repository says so in plain language in its own defensibility section: the thirteen layer decomposition is an obvious factoring of the problem, it is legible from a directory listing, and the orchestrator and run-state schema are not hard.

What takes longer is the gate rationales. Each one names the specific failure that gate exists to catch and shows the arithmetic behind its threshold. A real example, the rationale for G-L02-02, quoted in full:

60 is the smallest shortlist for which the quota's expected available count (21.15 under the declared design priors) exceeds L03's minimum of 3 eligible domains by more than 5x. Under 60, one adverse availability draw fails the entire run and forces a full regeneration.

That is not a line anyone writes before having a run fail on one adverse availability draw. The repository holds 395 of them.

What am I actually paying for, in one sentence?

A decided architecture and its written specification: what each of thirteen layers and seven sprint stages may read, must emit, and is judged against, plus 395 gates with the arithmetic behind every threshold, a working CLI, a working domain engine, four design prompts, and a documentation set that argues against its own product.

Why 199 dollars? Why not 49, or 999?

49 signals a template, and this isn't a template. 999 implies support, onboarding, and outcomes that are not being sold at any price here. 199 is priced for a competent operator who will read the contracts and run the thing themselves, and it is low enough that a fourteen day refund is a real offer rather than a formality.

Is the price per site, per year, or per seat?

None of those. One payment, one repository, yours permanently. No license server, no seat count, no renewal.

Do I get updates?

No promise of future work. A promise that cannot be verified at the moment of purchase does not belong in a product sold on verifiability. If updates happen you'll get them, and that sentence is not a commitment.

Section 4

About fit

Who is this for?

Two buyers.

The portfolio operator runs several small content or affiliate sites, already owns a registrar account, already has Search Console properties, can read a terminal, already knows what a canonical tag is. The job is reducing cost per experiment so they can take more shots per quarter. They understand that most sites fail, so a stop-loss rule reads as competence rather than as a disclaimer.

The agency or freelancer delivers sites for clients, is labour constrained, has margin pressure on the build phase. The job is producing a launch-ready, technically correct site in hours instead of days. Commercially that's the better fit, because they already have demand solved, so compressing build cost is worth paying for whether or not any given site ranks.

Who is this explicitly wrong for?

Someone who wants a website that gets customers and does not want to touch a terminal.

That buyer converts best on pages like this one and is the worst possible fit. They cannot complete the design layer, cannot debug a DNS record, will read "not indexed at day 30" as a defect rather than as a normal measurement, and will want a refund. The repository's own buyer analysis says exactly that, and it is repeated here so that person can save their money.

Also wrong for it: anyone wanting to produce hundreds of near-identical sites. The design deliberately caps pages per site and requires a non-reproducible element per site. It will fight you.

Do I need to be able to write Python?

Not to run it. The CLI is standard library only, Python 3.9 compatible, with no third-party dependencies to reconcile, and every assertion used by every gate is already implemented.

You do need to be comfortable reading gate output, editing JSON, and diagnosing a DNS record. If you want to add a layer or a gate, you will be writing Python, and the layer contract tells you exactly where it goes: assertions live in one library and a layer may not invent one without adding it there.

What exactly do I need before starting a site?

  • Python 3.9 or newer.
  • A GitHub account, plus git and the gh CLI, for deploy.
  • A Cloudflare account for DNS, edge TLS, caching, and the www-to-apex redirect.
  • A registrar account with a payment method on file, because layer 3 asks you to buy a .com inside a short decision window.
  • A Claude subscription, for running the four vendored prompts at claude.ai/design.
  • Model API credentials with a working balance, for the generative layers.
  • Keyword or SERP data credentials, if you want layers 1, 4 and 12 to hold real numbers rather than nulls.
  • A Google account with Search Console access, verified in advance rather than during layer 11.

How much will running one site cost me?

Unknown, and the repository refuses to guess. Its cost model is written as formulas with every price left as a variable, and the one worked example is labelled ILLUSTRATIVE ASSUMPTIONS ONLY with an instruction to replace every value with a measured one before making any pricing decision.

What the model does assert, robustly to the placeholders, is the ranking of the cost lines. The human design hour is almost certainly the largest, and it is the only line that does not fall with volume. Content generation tokens are second, dominated by page count. The domain is small in absolute terms but is spent before any validation, which makes it disproportionately painful across a portfolio of failures.

Section 5

About the design step

Why is there a manual step in an autonomous pipeline?

Because design taste is the thing a human has and a pipeline does not, and pretending otherwise produces sites that are technically correct and visibly generated.

Everything around layer 6 is automated specifically so your attention is spent only there. The layer's contribution is not doing the design, it is making the iteration converge: a weighted rubric with a numeric accept threshold and a per-dimension veto floor, revision requests required as diffs against named failing dimensions rather than restatements, and a minimum score gain per iteration below which the correct move is to restart from the brief rather than keep patching.

Without a stated threshold nobody ever chooses to restart, and design iteration without one is a random walk where you cannot tell whether iteration four is better than iteration two or merely different.

The acceptance gate is a manual attestation. It doesn't judge the design, it verifies that a human recorded a decision, a score, and a timestamp. That is the one place the pipeline cannot self-certify, so it refuses to try.

How long does the design step take?

The operating manual estimates 60 to 240 minutes of human attention, labels it the widest distribution in the run and the single largest cost line, and labels the estimate as an estimate. It is not a measurement. Nobody has stopwatched it across enough runs to have a median, and the productization document names that missing measurement as the single number that decides whether this is a software product or an agency service.

Are the four design prompts any good?

Judge them directly: they are 1,404 lines of vendored text in the repository and you can read all of them before deciding.

The strongest is the validator. It scores ten named dimensions, declares its own calibration (a 7 means better than 90 percent of what ships, an 8 could win an award, a 9 goes in the reviewer's portfolio, a 10 does not exist on first review), demands the top five failures each be located and paired with a specific fix rather than "make it better", and requires a re-score of the improved version before it will finish.

The sales page for this product was designed against that rubric, including its instruction not to default to a light gray background with a blue accent and Inter, which it names as the fingerprint of a template.

Section 6

About the mechanism

How is this different from a prompt chain?

Two rules, both visible in the terminal output on the sales page.

Nothing advances on a failed gate, including a gate that could not be evaluated. Results are PASS, FAIL, WARN, SKIP. SKIP means the gate could not be evaluated: a missing field, a non-numeric value where a number was required, an assertion that raised. A blocking gate returning SKIP fails the layer exactly as FAIL does. Treating "could not check" as "fine" is how automated pipelines ship broken output while reporting success. A gate that raises is caught and returns SKIP rather than crashing the run, so a broken gate degrades to "this halted the layer" and never to "this passed".

Unknown propagates as null, never as a flattering default. An unscreened domain does not score 1.0 on trademark clearance, it scores null, the weight vector renormalizes over the features that were actually measured, and the result carries a coverage fraction. A 0.80 at coverage 1.00 and a 0.80 at coverage 0.70 are visibly different objects, and the configuration sets a minimum evidence coverage floor independently of the minimum score.

The same rule governs content: the generation layer emits null for any statistic it could not source, and a gate fails on unsourced numerics. The pipeline is not permitted to invent a figure that reads as measured.

What happens if I re-run a layer that already passed?

Every downstream layer is reset to pending and its artifacts are moved into an invalidated directory rather than deleted.

The reset is enforced, not advised. An artifact derived from an input that has since changed is strictly worse than no artifact, because it carries the authority of having passed its gates while no longer being valid. Keeping it would mean your site's content was written against a keyword map that no longer exists.

Does the CLI call the model itself?

No, deliberately. awp prompt <layer> emits a self-contained execution brief (run context, declared upstream artifacts, the layer's agent prompt, and the verbatim gate list the output will be judged against). awp pass <layer> reads what the agent wrote and judges it.

Three consequences. The scaffolding is testable without spending tokens. Any agent runner can execute a layer, including a human. And failed attempts feed forward: attempt N+1 receives its predecessor's artifact plus the specific gate failures, so it corrects rather than restarts.

Agents are shown their gates on purpose. Writing to the test is the desired behaviour when the test encodes the actual quality bar.

Why is domain selection ordinary code instead of an agent?

Because it is string generation under a grammar, a local filter, a registry lookup, and a weighted sum. No judgement in that sequence is worth model tokens.

The cost asymmetry drives the design: every candidate killed locally is one rate-limited registry call not made. The canonical filter is not cosmetic polish, it is the cost control for the whole subsystem. In the run on the sales page it took 288 raw candidates down to 282 before any network call, and the nine filter stages each record a reason for every rejection.

What is the GEO sprint track, and why is it separate?

The GEO sprint track is the recurring post-launch growth cycle: 7 stages, 191 gates, run against a site that is already live and already measured.

It is separate from the layers for three reasons stated in its contract. It repeats, whereas a layer runs once and reaching passed is terminal. Its constraints are cycle-scoped rather than run-scoped (20 to 50 pages per sprint, a minimum 48 hours between deploys, live Search Console data required before any URL-level decision). And one constraint spans runs entirely: no two domains in your fleet may target the same primary keyword, which no per-run layer can enforce because a layer sees only its own run. That is why a fleet-wide keyword registry exists and why claiming a keyword is a gated action.

What is the stop-loss and why would I want one?

A stated time horizon and metric floor below which a site is abandoned rather than optimized forever. Without it, a productized pipeline burns its users' money indefinitely and calls it iteration.

It is deliberately hard to trigger: the abandon verdict requires all five criteria met simultaneously after at least 180 days, at least two completed monitoring cycles, and at least three changes applied and measured. By construction it is never a verdict about impatience.

The productization document argues the stop-loss should also be a pricing feature, because a vendor whose incentive is to bill optimisation retainers indefinitely will never ship one.

Section 7

About the transaction

What is the refund policy?

Fourteen days, full refund, no explanation required. Clone it, run ./verify.sh, run ./awp domains and ./awp gate, read the layer contracts. If the repository is not what the sales page describes, or you simply decide it is not for you, ask inside fourteen days.

Explicitly not covered: traffic, rankings, indexation, or revenue. No refund is available on the basis that a site did not rank, because no ranking is promised anywhere in this product. A guarantee attached to an outcome the seller does not control is not a guarantee, it is a liability being hidden until month three.

Is there support?

No commitment. Asynchronous and best effort, with no response time promised, and the price reflects that. If you need synchronous help, this is not the product and no amount of it will substitute.

Can you build the site for me?

Not at this price. The design step alone is one to four hours of human attention per site and it does not fall with volume, so done-for-you at 199 dollars would be a loss on the first site.

Can I resell it, or use it for client work?

Client work, yes. That is one of the two intended buyers.

Reselling the repository itself is a different question, and the repository is candid about the underlying reality: at this price point license terms are effectively unenforceable, and one buyer with a public repo ends the exclusivity. That is priced in. It is not an invitation.

Why does the sales page tell me so many reasons not to buy?

Because the alternative is selling a website generator while a buyer believes they are buying traffic, and that gap closes on schedule at roughly day 90 for every customer at once. The product's own risk register names expectation mismatch as a high-likelihood risk whose damage is mostly reputational, and notes that dissatisfied buyers of SEO-adjacent products are unusually vocal.

Setting the expectation on the sales page costs conversions. Setting it in the month-three support ticket costs the business.

The short version

What is not promised

  • NORankings. The pipeline produces zero off-site signal.
  • NOIndexation. Submitting through Search Console or IndexNow is a request, not an instruction.
  • NOTraffic on a timeline. No measured base rate exists, and any number given would be invented.
  • NOOutcome data. No end-to-end run has ever been completed, so there is no demo site and no measured median for anything.
  • NOSocial proof. No other buyers, no reviews, no case studies.
  • NOUpdates. A promise that cannot be verified at the moment of purchase does not belong in a product sold on verifiability.
  • NOSupport commitment. Asynchronous and best effort, with no response time promised.
  • NORemoval of the human hour. The design step is 60 to 240 minutes of human attention and it does not fall with volume.

Section 8

Reproducing every figure on this page

Run these in the repository root. If output disagrees with this page, the page is stale and the repository is right.

repository rootrun these
# the 14-check state report
./verify.sh

# layers, class, gate counts
./awp layers

# assertion names in use, and coverage
./awp assertions

# contract validation, all 13 layers
./awp doctor

# the live domain engine, then the gate evaluation quoted above
./awp domains demo --terms "espresso machine,coffee grinder,barista" --limit 40
./awp gate L03 demo

# the line counts quoted above
wc -l layers/*/LAYER.md layers/*/agent.md
wc -l sprints/*/LAYER.md sprints/*/agent.md
wc -l core/*.py docs/*.md layers/L06-design-handoff/prompts/*.md
Scroll the block horizontally on a narrow screen.

Live registry results vary by run, because availability changes. The raw and filtered candidate counts are deterministic and will not.