KiwopLabs:WhatWeTestBeforeWeRecommendIt

Reproducible experiments, measurements with method and dataset, and open engineering of our own systems. Primary source, not opinion: every piece links how it was measured and, when there is data, the dataset.

100 × 4 + 3/3 kiwop.com in the agentic browsing test
Scroll

What is Kiwop Labs?

Kiwop Labs is where we publish what we test before recommending it to a client: reproducible experiments (with the steps to repeat them), measurements with method and dataset (sample, criteria and date) and open engineering pieces about the systems we operate, starting with our own.

The reason is simple: in AI there is a lot of opinion and very little measurement. When a language model or a person looks for an answer, they cite whoever provides the original data and explains how it was obtained. So there are no figures without a method here, and if a data point contradicts our thesis, we change the thesis (it already happened with the ecommerce study: the pilot knocked down the starting hypothesis and we published it anyway).

Labs in Numbers

Only figures already published on this site, with their source.

100 × 4 Lighthouse on kiwop.com

Performance, accessibility, best practices and SEO, in the agentic browsing test

3 / 3 Agentic tasks completed

An agent navigates, finds and completes the three tasks of the test

300 Stores measured

"State of AI in Spanish ecommerce 2026" study, double pass

7 Languages

Every piece, in the seven languages of the site

Experiments and Repros

Each one with the steps to repeat it.

01

Agentic browsing test

We check whether an AI agent can navigate your website and complete real tasks, with the Lighthouse audits and three agentic tasks. kiwop.com scores 100 in all four categories and completes 3 out of 3. Run the test on your site.

02

WebMCP and the Chrome 150 crash

We enabled WebMCP in production and Chrome killed the tab on the first click. We bisected the trigger (router + imperative tools + declarative tool at the same time) and published the repro for Chromium. The experiment and the public repro.

03

MCP 2026-07-28: the stateless protocol

We read the fifth revision of MCP the day it shipped and explained what really changes (sessions and handshake are gone, routing headers), what to migrate urgently and what not to. Read the analysis.

04

RAG, fine-tuning or MCP: when to use each

Three techniques that get confused every day, with the cases where we have used each one in production and the criteria to choose. Read the guide.

05

How we use AI to code at an agency

No hype: how we work with Claude Code day to day, what we review by hand, where it fails and what we have stopped doing. Read how we work.

06

Nexo: multi-tenant SaaS architecture on Laravel + Inertia

The platform we run Kiwop on, opened up: workspace isolation, role-based permissions, the agent layer with human approval and what we learned. The architecture and the case with real metrics.

Data and Measurements

Sample, criteria and date, always published with the figure.

State of AI in Spanish ecommerce 2026

300 stores selected with a public rule (sector × size), measured with a double pass (identified HTTP and a real browser) in September 2026: AI crawlers in robots.txt, llms.txt, Product schema, chatbots and search. The pilot knocked down the initial hypothesis and we changed the thesis. Dataset and method are published with the report.

Publication: September 2026

Nexo: metrics from a platform with agents in production

Active agents, logged executions, tool calls and tokens processed, extracted through a read-only query against the production database, with the date and method on the case page itself. See the metrics.

27 agents in production

La Salve: an AI chatbot with guardrails, in production

How the assistant of a real ecommerce is built: what the model decides, what runs through deterministic logic and what its limits are. Operating metrics are published with period and method once the client validates them. See the case.

Architecture and guardrails published

kiwop.com as a test bench

Our own website is the first lab: the indexing of the AI pages Google had left unindexed for months, recovered the same day by cutting off a domain mirror, with the census and the method. See the case.

10 of 10 URLs indexed the same day

Open Tools

Free, no sign-up, in seven languages.

01

Agentic browsing test

Can an AI agent navigate your website and complete a task? Lighthouse in all four categories and three agentic tasks, with the result explained. Run the test.

02

EU AI Act self-assessment

A short questionnaire to find out whether your AI system falls under the European regulation, in which risk category and which obligations apply to you, with verified dates. Take the self-assessment.

03

How much does an AI project cost?

Investment ranges by type of AI project, with what each tier includes and what makes it more expensive. Calculate the range.

04

AI skills test

A short test to find out where your team stands with AI and which Kiwop Academy course fits. Take the test.

How We Work in Labs

The rules that make a piece citable.

Method before result: sample, criteria and date are published with the figure, and in the ecommerce study the protocol was frozen before measuring
Zero invented figures: if it has not been measured, it is not published; if the client has not validated it, it does not carry their name
Repro before opinion: an experiment includes the steps to repeat it, and a bug includes the repro
Corrected in public: when the data contradicts the thesis, the thesis changes and it stays on the record
At home first: kiwop.com and Nexo are the first test bench for everything we recommend
Seven languages: every piece is published in the seven languages of the site, with its own URL

How to Propose an Experiment

If you have a question that deserves a measurement, we run it.

01

You tell us the question

What you want to know and why it matters for your business: whether your site is ready for agents, whether a RAG improves your customer service, how much of your traffic already comes from a language model.

02

We design the method and publish it before measuring

Sample, criteria, tools and what would count as a result. Published before measuring so nobody can move the goalposts afterwards.

03

We measure and publish, whatever the data says

The result ships with the method and the dataset. If it contradicts what we expected, we publish it anyway: that is what makes the next piece believable.

At Home First, Then at the Client

Everything we publish in Labs has been tested first on our own systems: Nexo, the platform we run Kiwop on, runs 27 AI agents in production with human approval and traceability; and kiwop.com is the site where we validate performance, accessibility, agentic browsing and WebMCP before touching anyone else's. When something breaks, like the Chrome 150 crash, it gets documented and the repro gets published.

6 Primary-source pieces published
4 Open tools
0 Figures without a method

Questions About Kiwop Labs

What people ask us about this section.

What is Kiwop Labs?

The section of kiwop.com where we publish primary source: reproducible experiments, measurements with method and dataset, and open engineering of the systems we operate. It is not the blog: only what has been measured or built gets in, with the steps to repeat it.

Can I reproduce the experiments?

Yes, that is the condition for publishing them. Every experiment includes the steps, the tools and the versions. The WebMCP repro is published exactly as we sent it to Chromium, and you can run the agentic browsing test on your own website.

Do you publish the data?

When there is data, yes: sample, selection criteria and date go with the result, and the dataset is published with the report (for the ecommerce study, the 300 domains and the selection rule). What we do not publish is client data without their validation or proprietary figures from third parties.

Why does an agency publish this?

Because it is the most honest way to show we know how to do it: original data gets cited, opinion does not. A language model answering "which agency knows about AI agents" links to whoever published the measurement, not to whoever claims it. And it forces us to test before we recommend.

Can I propose an experiment?

Yes. Tell us the question and why it matters for your business. If it deserves a measurement, we design the method, publish it before measuring and publish the result whatever it says. Write to us.

How often do you publish?

When something has been measured, not on a calendar. The goal from September 2026 is one or two pieces a month, always with a method and in the seven languages of the site.

Does this replace your services?

No, it supports them. What we learn in Labs is what we apply in AI agents, enterprise RAG, GEO and AI audits for clients. If you are looking for a project, start there.

Got a question that deserves an experiment?

Tell us. If it deserves a measurement, we publish the method before measuring and the result whatever it says.

Propose an experiment
No commitment Response in 24h Custom proposal
Last updated: September 2026

Initial technical
consultation.

AI, security and performance. Diagnosis with phased proposal.

NDA available
Response <24h
Phased proposal

Your first meeting is with a Solutions Architect, not a salesperson.

Request diagnosis