VLIRTZ

Built for the Netherlands, from Stockholm

AI agent development in Amsterdam

We build AI agents that run one real workflow end to end. Something working on your own data in the first week or two, published price bands, and a human on every action that moves money.

Dutch operators want to see it working before they will discuss a roadmap, and they ask the price in the first ten minutes. Both instincts are right, so we publish the bands and build the prototype first.

Prototype in week one, roadmap last

Dutch clients have been unusually direct with us about wanting to see something work before discussing strategy, and it has shaped how we sequence a build. The first deliverable runs on your data, including the messy records. Not a demo on a curated sample.

The reason is not showmanship. A prototype on real data is the fastest way to find out that thirty percent of your source records are missing a field the workflow depends on, or that the process everyone described has three undocumented exceptions. That discovery is worth more than any strategy document, and it arrives in week one rather than month three.

Only after that do we close the agent loop, and only after that discuss a roadmap. Roadmaps written before anyone has touched the data are fiction, and Amsterdam operators seem to spot that faster than most markets.

What it costs, and what moves the number

We publish bands rather than making you extract them over three calls. The vendors Dutch buyers actually find in search publish their numbers, and treating price as a secret mostly wastes the time of people who were never going to fit the budget anyway.

What moves a quote inside a band: how many systems the agent has to touch, how usable your data already is before anyone cleans it, how expensive a wrong action would be, and whether your team or ours operates it afterwards. A range quoted without those drivers is a guess dressed as an estimate.

We also separate build cost from running cost, because they get conflated and the second one causes the surprises. Model and API usage is usually the smallest line. Maintenance is the real recurring cost, since models change, your source systems change, and an unmaintained agent degrades quietly rather than failing loudly.

WBSO, and why the documentation matters

Development work with genuine technical uncertainty often qualifies for the Dutch WBSO R&D tax credit, which materially lowers the net cost of a build. Agent work frequently involves exactly that kind of uncertainty, particularly around retrieval quality and evaluation on messy real-world data.

We structure our engineering documentation so the work is legible to a WBSO application: what was uncertain at the outset, what was attempted, what was measured, and what was concluded. That is documentation worth having regardless, because it is also what lets someone change the system safely a year later.

To be clear about the boundary: we are not tax advisors and we do not assess eligibility or file on your behalf. Your own advisor does that. What we can do is make sure the engineering record supports the claim rather than working against it.

Use cases

What we are asked to build in Amsterdam

Real workflow shapes from this market, with the point where a human stays in the loop stated for each one.

Fintech and payments

Reconciliation and dispute handling

High-volume ledger mismatches and chargeback cases where resolution means checking several systems. The agent assembles context and proposes the resolution. Anything moving funds stays behind human approval, because a wrong action here costs money directly.

Fintech and payments

Merchant onboarding checks

Register lookups, document collection and the sanity checks a compliance analyst runs before approval. The agent gathers and cross-references, then hands over a complete file with discrepancies flagged.

Logistics and supply chain

Freight document extraction

Customs paperwork, manifests and delivery exceptions across formats never designed to be machine-readable. Retrieval quality decides the outcome here rather than reasoning ability, and we scope it on that basis.

Scale-ups

Support and sales-ops triage

The classic scale-up bottleneck: a small team drowning in requests that each need a lookup before they can be answered. Narrow, measurable, painful enough that people actually adopt the result, which makes it a good first project.

Marketplaces

Listing and content moderation queues

Policy checks at volume where most cases are clear and the edge cases matter. The agent handles the clear ones and routes the rest with the policy citation attached. Anything affecting a user's access to the platform keeps a human in the loop.

How we build

A Amsterdam agent build, week by week

Two to four weeks from kickoff to handover. The order matters more than the tooling: measure first, prototype on real data second, close the loop third.

  1. 01

    Watch the workflow being done

    2 to 4 days

    We sit with the people who run the process today, and we measure it: how many cases, how long each takes, where they stall, and which exceptions actually recur. Most projects that fail do so because this step was skipped and the brief described the process as management believes it works rather than as it runs.

    You end up with: A measured baseline you can hold the finished system against, and a written list of the exceptions nobody had documented.

  2. 02

    Build the thin version on your real data

    3 to 5 days

    Not a demo on a curated sample. Your records, including the ones with missing fields and inconsistent formatting. This is where you discover that a third of the source rows lack something the workflow depends on, and it is much better to discover that in week one than in month three.

    You end up with: A narrow tool running on production-shaped data, and an honest assessment of whether the rest is worth building.

  3. 03

    Close the agent loop

    1 to 2 weeks

    Now the agent plans across steps, calls the tools it needs, and handles the cases the thin version could not. Human review gates go on every action that is expensive to undo. We build the evaluation set from your real cases at the same time, including the failures, because an agent with no evaluation set is an agent nobody can safely change later.

    You end up with: A working agent, an evaluation suite built from your own cases, and audit logging on every tool call.

  4. 04

    Hand it over properly

    2 to 4 days

    A runbook, a training session with the people who will operate it, and a documented path for what to do when it breaks. We do not make handover deliberately incomplete to keep you dependent on us. If you want us to keep operating it, that is a separate retainer you choose, not a trap you fall into.

    You end up with: Runbook, handover session, and the code and configuration in your own repository.

How we build

Positions we hold on every build

These are decisions we make the same way every time, because each one is a reason agent projects fail when it goes the other way.

Bounded autonomy by default

An agent starts read-only and draft-first. Actions that are expensive or awkward to reverse stay behind a human approval gate, permanently if that is the right answer. Full autonomy is something a system earns by demonstrating accuracy on your evaluation set, not a launch feature.

An evaluation set from your real cases

Built from your actual records, including the ones the agent gets wrong. Without it, nobody can safely change a prompt or swap a model six months later, which is how working systems quietly rot.

Every tool call logged

Each action, its inputs, and the decision behind it are recorded. This is what makes an incident investigable, and under FINMA, NIS2 or medical-device rules it is a requirement rather than a nicety.

Retrieval quality over model size

Most disappointing agents are not under-powered, they are under-informed. Getting the right context in front of the model reliably matters more than which model it is, and it is where the engineering effort usually belongs.

Your repository, your infrastructure

Code and configuration live in your repository and run on infrastructure you control. There is no VLIRTZ platform you have to keep paying for to keep your own workflow running.

One workflow before three

We decline company-wide assistant scopes. A single workflow, measured and shipped, tells you more about whether this approach works for you than any roadmap, and it is recoverable if the answer is no.

Pricing

What an AI agent costs in Amsterdam

We quote in EUR for this market. Where you land depends on how many systems the agent touches, how usable your data already is, and how expensive a wrong action would be. Ask and you get a range on the first call, not the third.

Agent feasibility review

On request

We measure the workflow, assess whether your data supports it, and tell you whether an agent is the right answer. Includes a scoped build proposal.

Timeline: 1 to 2 weeks

Scoped agent build

On request

One workflow end to end, integrated with your tools, human review gates, evaluation set, audit logging, runbook and handover.

Timeline: 2 to 4 weeks

Additional workflow

On request

A second or third agent reusing the orchestration, retrieval and evaluation harness from the first.

Timeline: 2 to 3 weeks each

Sustain retainer

On request

Monitoring, drift checks, model and prompt updates, and a defined response time on failures.

Timeline: Rolling

Straight answers

What an AI agent will not do for you

Every one of these has ended a project somewhere. We would rather raise them before you sign than explain them in month two.

It will not fix a process nobody has agreed on
If two departments genuinely disagree about how a case should be handled, an agent forces that disagreement into the open rather than resolving it. That is useful, but it is a management outcome, not a technical one.
It will not rescue unusable source data
Retrieval over clean, structured records is straightforward. Retrieval over scanned documents, three competing sources of truth, and a field that has been wrong since a migration is where budgets disappear. Sometimes the honest recommendation is a data project first.
It will not be right every time
The question is never whether it makes mistakes, it is whether the mistakes are caught before they cost anything. That is what the review gates and the evaluation set are for, and it is why we measure the baseline first.
It will not maintain itself
Models change and your source systems change. An unmaintained agent degrades quietly rather than failing loudly, which is worse. Budget for maintenance or plan to retire it.

More about working with us in Amsterdam

This page covers how we build agents. The Amsterdam market page covers the rest: the regulators that shape a project there, the sectors we see most, and when we are the wrong partner.

AI Software Agency for Amsterdam Companies

FAQ

AI agent development in Amsterdam: common questions

How much does AI agent development cost in Amsterdam?
We publish bands rather than making you ask three times. A scoped single-workflow build is the usual entry point, with a shorter feasibility review ahead of it if the use case is not settled. What moves the number is systems touched, data usability, cost of a wrong action, and who operates it afterwards. Full breakdown on the pricing page.
How quickly will we see something working?
The first deliverable runs on your own data, including the messy records, usually inside the first week or two. That is deliberate. A prototype on real data surfaces the missing fields and undocumented exceptions that would otherwise derail the project in month three.
How long does the whole build take?
Two to four weeks from kickoff to handover for a scoped single-workflow agent. Anything quoted at under a week is a demo rather than a system, and anything open-ended should worry you.
Does our agent project qualify for WBSO?
Development with genuine technical uncertainty often does, and agent work frequently involves exactly that around retrieval quality and evaluation on messy data. We structure the engineering documentation so the work is legible to an application, but your own tax advisor assesses eligibility and files it. We are not tax advisors.
Will the agent act on its own?
It starts read-only and draft-first, and anything expensive to undo stays behind human approval. For Dutch fintech work in particular, anything that moves funds keeps a human gate permanently. Autonomy is earned on your evaluation set over real volume, not switched on at launch.
What about the AP and automated decisions?
It matters if your workflow ranks, scores or filters people. The Dutch authority has been among Europe's more assertive on automated risk scoring, so anything touching hiring, credit or access to services needs deliberate classification and a documented human review step rather than an assumption that it is low risk.
Which frameworks and models do you use?
We are not tied to one. Most builds use an orchestration layer such as LangGraph with a hosted model and a vector store for retrieval, but the choice follows your data residency and latency constraints. Retrieval quality affects the outcome more than which model sits behind it, which is where the engineering effort actually belongs.
What do we own at the end?
Code and configuration in your own repository, running on infrastructure you control, plus the evaluation set, the audit logging and a runbook. There is no platform of ours you have to keep paying for.
Do you have an office in Amsterdam?
No. We are headquartered in Stockholms lan and work with Dutch clients remotely, travelling for kickoff and key milestones. There is no time difference at all, so the working day overlaps completely.

This page was last reviewed on .

By market

Agent development in other markets

Each market page covers the workflows we are actually asked to build there, the regulators that shape the design, and pricing in the local currency.