Kiwop Labs: What We Test Before We Recommend It

Reproducible experiments, measurements with method and dataset, and open engineering of our own systems. Primary source, not opinion: every piece links how it was measured and, when there is data, the dataset.

  • 100 × 4 + 3/3 kiwop.com in the agentic browsing test
Scroll

Definition

What is Kiwop Labs?

Kiwop Labs is where we publish what we test before recommending it to a client: reproducible experiments (with the steps to repeat them), measurements with method and dataset (sample, criteria and date) and open engineering pieces about the systems we operate, starting with our own.

The reason is simple: in AI there is a lot of opinion and very little measurement. When a language model or a person looks for an answer, they cite whoever provides the original data and explains how it was obtained. So there are no figures without a method here, and if a data point contradicts our thesis, we change the thesis (it already happened with the ecommerce study: the pilot knocked down the starting hypothesis and we published it anyway).

In numbers

Labs in Numbers

Only figures already published on this site, with their source.

  • 100 × 4 Lighthouse on kiwop.com Performance, accessibility, best practices and SEO, in the agentic browsing test
  • 3 / 3 Agentic tasks completed An agent navigates, finds and completes the three tasks of the test
  • 300 Stores measured "State of AI in Spanish ecommerce 2026" study, double pass
  • 7 Languages Every piece, in the seven languages of the site

Key points

Experiments and Repros

Each one with the steps to repeat it.

  1. 01

    Agentic browsing test

    We check whether an AI agent can navigate your website and complete real tasks, with the Lighthouse audits and three agentic tasks. kiwop.com scores 100 in all four categories and completes 3 out of 3. Run the test on your site.

  2. 02

    WebMCP and the Chrome 150 crash

    We enabled WebMCP in production and Chrome killed the tab on the first click. We bisected the trigger (router + imperative tools + declarative tool at the same time) and published the repro for Chromium. The experiment and the public repro.

  3. 03

    MCP 2026-07-28: the stateless protocol

    We read the fifth revision of MCP the day it shipped and explained what really changes (sessions and handshake are gone, routing headers), what to migrate urgently and what not to. Read the analysis.

  4. 04

    RAG, fine-tuning or MCP: when to use each

    Three techniques that get confused every day, with the cases where we have used each one in production and the criteria to choose. Read the guide.

  5. 05

    How we use AI to code at an agency

    No hype: how we work with Claude Code day to day, what we review by hand, where it fails and what we have stopped doing. Read how we work.

  6. 06

    Nexo: multi-tenant SaaS architecture on Laravel + Inertia

    The platform we run Kiwop on, opened up: workspace isolation, role-based permissions, the agent layer with human approval and what we learned. The architecture and the case with real metrics.

Use cases

Data and Measurements

Sample, criteria and date, always published with the figure.

  • GEO Baseline: which agencies LLMs cite in Spain

    Every month we send 40 buying questions (15 from the marketingdirecto ranking, reproduced verbatim, 15 of our own about AI agencies and 10 with what people really ask assistants) to search-enabled assistants and publish which agencies they recommend, which domains they cite and where Kiwop appears, with the full answers. See the series.

    Monthly series, open dataset

  • What Spain asks AI, versus Google

    Every month, how many queries AI assistants receive versus Google for 60 terms about agencies, prices and AI for business. Price questions already go to AI first; local searches stay on Google. See the series.

    Monthly, open dataset

  • State of AI in Spanish ecommerce 2026

    300 stores selected with a public rule (sector × size), measured with a double pass (identified HTTP and a real browser) in September 2026: AI crawlers in robots.txt, llms.txt, Product schema, chatbots and search. The pilot knocked down the initial hypothesis and we changed the thesis. Dataset and method are published with the report. Read the report.

    Published: 300 stores, open dataset

  • Evidence-based ranking of AI agencies

    Every month we count in how many ChatGPT, Claude, Gemini and Perplexity answers each Spanish AI, GEO and marketing agency appears, with the full answers and the domains they cite. Kiwop lands where it lands, and we say so. See the ranking.

    Monthly, recalculated automatically

  • Published prices of digital services in Spain

    Every quarter we go back to the pricing pages of Spanish agencies and freelancers and check that the price they publish is still there: website, online store, SEO, maintenance, development hour, AI, chatbots, GEO and Ads. Only prices published by the provider itself, with literal text and URL; Kiwop is on the list and counts like everyone else. See the observatory.

    Quarterly, open dataset

  • AI agents in production: Nexo's telemetry, every quarter

    Agent runs and failures, proposals signed or discarded by a person, PRs delivered by the autonomous worker, classified email and cost via API and via subscription. Read-only aggregates from the platform Kiwop runs on, with no names or texts. See the series.

    Quarterly, open SQL query

  • Nexo: metrics from a platform with agents in production

    Active agents, logged executions, tool calls and tokens processed, extracted through a read-only query against the production database, with the date and method on the case page itself. See the metrics.

    27 agents in production

  • La Salve: an AI chatbot with guardrails, in production

    How the assistant of a real ecommerce is built: what the model decides, what runs through deterministic logic and what its limits are. Operating metrics are published with period and method once the client validates them. See the case.

    Architecture and guardrails published

  • kiwop.com as a test bench

    Our own website is the first lab: the indexing of the AI pages Google had left unindexed for months, recovered the same day by cutting off a domain mirror, with the census and the method. See the case.

    10 of 10 URLs indexed the same day

Key points

Open Tools

Free, no sign-up, in seven languages.

  1. 01

    Agentic browsing test

    Can an AI agent navigate your website and complete a task? Lighthouse in all four categories and three agentic tasks, with the result explained. Run the test.

  2. 02

    EU AI Act self-assessment

    A short questionnaire to find out whether your AI system falls under the European regulation, in which risk category and which obligations apply to you, with verified dates. Take the self-assessment.

  3. 03

    How much does an AI project cost?

    Investment ranges by type of AI project, with what each tier includes and what makes it more expensive. Calculate the range.

  4. 04

    AI skills test

    A short test to find out where your team stands with AI and which Kiwop Academy course fits. Take the test.

  5. 05

    Is your store ready for AI?

    Type the domain of an online store and in a minute you get a public report with what an AI assistant sees: whether it can get in, whether there is a real llms.txt and whether the product page exposes price and stock. Same method as the 300-store study. Check a store.

What's included

How We Work in Labs

The rules that make a piece citable.

  • Method before result: sample, criteria and date are published with the figure, and in the ecommerce study the protocol was frozen before measuring
  • Zero invented figures: if it has not been measured, it is not published; if the client has not validated it, it does not carry their name
  • Repro before opinion: an experiment includes the steps to repeat it, and a bug includes the repro
  • Corrected in public: when the data contradicts the thesis, the thesis changes and it stays on the record
  • At home first: kiwop.com and Nexo are the first test bench for everything we recommend
  • Seven languages: every piece is published in the seven languages of the site, with its own URL

How we work

How to Propose an Experiment

If you have a question that deserves a measurement, we run it.

  1. 01

    You tell us the question

    What you want to know and why it matters for your business: whether your site is ready for agents, whether a RAG improves your customer service, how much of your traffic already comes from a language model.

    One call
  2. 02

    We design the method and publish it before measuring

    Sample, criteria, tools and what would count as a result. Published before measuring so nobody can move the goalposts afterwards.

    One week
  3. 03

    We measure and publish, whatever the data says

    The result ships with the method and the dataset. If it contradicts what we expected, we publish it anyway: that is what makes the next piece believable.

    Depends on the experiment

The proof

At Home First, Then at the Client

Everything we publish in Labs has been tested first on our own systems: Nexo, the platform we run Kiwop on, runs 27 AI agents in production with human approval and traceability; and kiwop.com is the site where we validate performance, accessibility, agentic browsing and WebMCP before touching anyone else's. When something breaks, like the Chrome 150 crash, it gets documented and the repro gets published.

  • 8 Primary-source pieces published
  • 5 Open tools
  • 0 Figures without a method

FAQ

Questions About Kiwop Labs

What people ask us about this section.

What is Kiwop Labs?

The section of kiwop.com where we publish primary source: reproducible experiments, measurements with method and dataset, and open engineering of the systems we operate. It is not the blog: only what has been measured or built gets in, with the steps to repeat it.

Can I reproduce the experiments?

Yes, that is the condition for publishing them. Every experiment includes the steps, the tools and the versions. The WebMCP repro is published exactly as we sent it to Chromium, and you can run the agentic browsing test on your own website.

Do you publish the data?

When there is data, yes: sample, selection criteria and date go with the result, and the dataset is published with the report (for the ecommerce study, the 300 domains and the selection rule). What we do not publish is client data without their validation or proprietary figures from third parties.

Why does an agency publish this?

Because it is the most honest way to show we know how to do it: original data gets cited, opinion does not. A language model answering "which agency knows about AI agents" links to whoever published the measurement, not to whoever claims it. And it forces us to test before we recommend.

Can I propose an experiment?

Yes. Tell us the question and why it matters for your business. If it deserves a measurement, we design the method, publish it before measuring and publish the result whatever it says. Write to us.

How often do you publish?

When something has been measured, not on a calendar. The goal from September 2026 is one or two pieces a month, always with a method and in the seven languages of the site.

Does this replace your services?

No, it supports them. What we learn in Labs is what we apply in AI agents, enterprise RAG, GEO and AI audits for clients. If you are looking for a project, start there.

Next step

Got a question that deserves an experiment?

Tell us. If it deserves a measurement, we publish the method before measuring and the result whatever it says.

  • No commitment
  • Response in 24h
  • Custom proposal
Last updated: September 2026

Let's talk.

Initial technical consultation

AI, security and performance. Diagnosis with phased proposal.

  • NDA available
  • Response <24h
  • Phased proposal

Your first meeting is with a Solutions Architect, not a salesperson.

Request diagnosis