KiwopLabs:WhatWeTestBeforeWeRecommendIt
Reproducible experiments, measurements with method and dataset, and open engineering of our own systems. Primary source, not opinion: every piece links how it was measured and, when there is data, the dataset.
What is Kiwop Labs?
Kiwop Labs is where we publish what we test before recommending it to a client: reproducible experiments (with the steps to repeat them), measurements with method and dataset (sample, criteria and date) and open engineering pieces about the systems we operate, starting with our own.
Labs in Numbers
Only figures already published on this site, with their source.
Performance, accessibility, best practices and SEO, in the agentic browsing test
An agent navigates, finds and completes the three tasks of the test
"State of AI in Spanish ecommerce 2026" study, double pass
Every piece, in the seven languages of the site
Experiments and Repros
Each one with the steps to repeat it.
Agentic browsing test
We check whether an AI agent can navigate your website and complete real tasks, with the Lighthouse audits and three agentic tasks. kiwop.com scores 100 in all four categories and completes 3 out of 3. Run the test on your site.
WebMCP and the Chrome 150 crash
We enabled WebMCP in production and Chrome killed the tab on the first click. We bisected the trigger (router + imperative tools + declarative tool at the same time) and published the repro for Chromium. The experiment and the public repro.
MCP 2026-07-28: the stateless protocol
We read the fifth revision of MCP the day it shipped and explained what really changes (sessions and handshake are gone, routing headers), what to migrate urgently and what not to. Read the analysis.
RAG, fine-tuning or MCP: when to use each
Three techniques that get confused every day, with the cases where we have used each one in production and the criteria to choose. Read the guide.
How we use AI to code at an agency
No hype: how we work with Claude Code day to day, what we review by hand, where it fails and what we have stopped doing. Read how we work.
Nexo: multi-tenant SaaS architecture on Laravel + Inertia
The platform we run Kiwop on, opened up: workspace isolation, role-based permissions, the agent layer with human approval and what we learned. The architecture and the case with real metrics.
Data and Measurements
Sample, criteria and date, always published with the figure.
State of AI in Spanish ecommerce 2026
300 stores selected with a public rule (sector × size), measured with a double pass (identified HTTP and a real browser) in September 2026: AI crawlers in robots.txt, llms.txt, Product schema, chatbots and search. The pilot knocked down the initial hypothesis and we changed the thesis. Dataset and method are published with the report.
Publication: September 2026
Nexo: metrics from a platform with agents in production
Active agents, logged executions, tool calls and tokens processed, extracted through a read-only query against the production database, with the date and method on the case page itself. See the metrics.
27 agents in production
La Salve: an AI chatbot with guardrails, in production
How the assistant of a real ecommerce is built: what the model decides, what runs through deterministic logic and what its limits are. Operating metrics are published with period and method once the client validates them. See the case.
Architecture and guardrails published
kiwop.com as a test bench
Our own website is the first lab: the indexing of the AI pages Google had left unindexed for months, recovered the same day by cutting off a domain mirror, with the census and the method. See the case.
10 of 10 URLs indexed the same day
Open Tools
Free, no sign-up, in seven languages.
Agentic browsing test
Can an AI agent navigate your website and complete a task? Lighthouse in all four categories and three agentic tasks, with the result explained. Run the test.
EU AI Act self-assessment
A short questionnaire to find out whether your AI system falls under the European regulation, in which risk category and which obligations apply to you, with verified dates. Take the self-assessment.
How much does an AI project cost?
Investment ranges by type of AI project, with what each tier includes and what makes it more expensive. Calculate the range.
AI skills test
A short test to find out where your team stands with AI and which Kiwop Academy course fits. Take the test.
How We Work in Labs
The rules that make a piece citable.
How to Propose an Experiment
If you have a question that deserves a measurement, we run it.
You tell us the question
What you want to know and why it matters for your business: whether your site is ready for agents, whether a RAG improves your customer service, how much of your traffic already comes from a language model.
We design the method and publish it before measuring
Sample, criteria, tools and what would count as a result. Published before measuring so nobody can move the goalposts afterwards.
We measure and publish, whatever the data says
The result ships with the method and the dataset. If it contradicts what we expected, we publish it anyway: that is what makes the next piece believable.
At Home First, Then at the Client
Everything we publish in Labs has been tested first on our own systems: Nexo, the platform we run Kiwop on, runs 27 AI agents in production with human approval and traceability; and kiwop.com is the site where we validate performance, accessibility, agentic browsing and WebMCP before touching anyone else's. When something breaks, like the Chrome 150 crash, it gets documented and the repro gets published.
Questions About Kiwop Labs
What people ask us about this section.
What is Kiwop Labs?
The section of kiwop.com where we publish primary source: reproducible experiments, measurements with method and dataset, and open engineering of the systems we operate. It is not the blog: only what has been measured or built gets in, with the steps to repeat it.
Can I reproduce the experiments?
Yes, that is the condition for publishing them. Every experiment includes the steps, the tools and the versions. The WebMCP repro is published exactly as we sent it to Chromium, and you can run the agentic browsing test on your own website.
Do you publish the data?
When there is data, yes: sample, selection criteria and date go with the result, and the dataset is published with the report (for the ecommerce study, the 300 domains and the selection rule). What we do not publish is client data without their validation or proprietary figures from third parties.
Why does an agency publish this?
Because it is the most honest way to show we know how to do it: original data gets cited, opinion does not. A language model answering "which agency knows about AI agents" links to whoever published the measurement, not to whoever claims it. And it forces us to test before we recommend.
Can I propose an experiment?
Yes. Tell us the question and why it matters for your business. If it deserves a measurement, we design the method, publish it before measuring and publish the result whatever it says. Write to us.
How often do you publish?
When something has been measured, not on a calendar. The goal from September 2026 is one or two pieces a month, always with a method and in the seven languages of the site.
Does this replace your services?
No, it supports them. What we learn in Labs is what we apply in AI agents, enterprise RAG, GEO and AI audits for clients. If you are looking for a project, start there.
Got a question that deserves an experiment?
Tell us. If it deserves a measurement, we publish the method before measuring and the result whatever it says.
Propose an experiment Initial technical
consultation.
AI, security and performance. Diagnosis with phased proposal.
Your first meeting is with a Solutions Architect, not a salesperson.
Request diagnosis