VLIRTZ

Built from Stockholms lan

AI agent development in Stockholm

We build AI agents that run one real workflow end to end, in the tools your team already uses, with a human on every decision that is expensive to undo. Based in Stockholms lan, so the workshops happen in your office.

Stockholm buyers have usually seen an agent demo that worked on curated data and failed on theirs. We start by measuring your workflow as it actually runs, then build the thin version on your messiest records first.

Why we build the thin version before the agent

The first working thing we hand a Stockholm client is usually not an agent. It is a narrow tool that does one step of the workflow on their real data, including the records with missing fields and inconsistent formatting.

This is not caution, it is speed. A prototype on real data is the fastest available way to find out that a third of the source rows are missing something the process depends on, or that the workflow everyone described has three undocumented exceptions. Discovering that in the first week changes what we build. Discovering it in month three means rebuilding.

It also gives you an early, cheap exit. If the thin version shows the data is not usable, you have spent days rather than a full build budget, and the honest recommendation is a data project first. We have made that recommendation and we would rather make it than deliver something that quietly underperforms.

Where the human stays, and why that is a design decision

Every agent we build starts read-only and draft-first. It reads the sources of truth, proposes an action, and a person approves it. Actions that are expensive or awkward to reverse stay behind that gate permanently if that is the right answer for the workflow.

Autonomy is something a system earns by demonstrating accuracy on your own evaluation set, over real volume, not something switched on at launch because it demos better. In Swedish fintech and payments work this is rarely controversial, because the cost of a wrong action is obvious and Finansinspektionen's expectations around reversibility point the same way.

The practical consequence is that a draft-and-approve agent is far easier to make genuinely reliable than an autonomous one, and it captures most of the time saving. The bottleneck in these workflows is almost always the lookup and assembly, not the final click.

What being in Stockholms lan changes about the build

Agent projects fail on context far more often than on capability. The engineer does not know that the finance team ignores one field because it has been wrong since a 2019 migration, or that the exception queue is actually triaged by someone in a different department. That knowledge does not survive a written brief.

Because we are based in Stockholms lan, the observation session where that surfaces is the default rather than a line item cut for budget. Kickoff, mid-build review and handover training happen in your office. Workshops can run in Swedish; delivery documentation stays in English by default because it tends to outlive team changes.

For Stockholm clients this is the concrete argument for hiring locally over hiring remotely, and it is the one part of our offer that a remote competitor genuinely cannot match.

Use cases

What we are asked to build in Stockholm

Real workflow shapes from this market, with the point where a human stays in the loop stated for each one.

Fintech and payments

Reconciliation exception handling

High-volume ledgers throw a steady stream of mismatches that a person resolves by checking three systems. The agent assembles that context, proposes the resolution, and applies it where the pattern is unambiguous. Anything that moves funds stays behind human approval.

Fintech and payments

Merchant and customer onboarding checks

Document collection, register lookups, and the sanity checks a compliance analyst performs before approving an account. The agent gathers and cross-references, then hands a complete file to the analyst with the discrepancies flagged rather than buried.

Consumer software and gaming

Support ticket triage with lookup

Most tickets need someone to check an account state, an order, or a log before answering. The agent does that lookup, drafts the reply, and routes anything unusual to a human with the evidence already attached.

Logistics and e-commerce

Logistics exception triage

Shipment and inventory exceptions where one experienced person is both the bottleneck and the single point of failure. The agent handles the cases where their judgement is genuinely a rule, and escalates the rest with the context assembled.

Professional services

Inbound lead qualification

Enrich an inbound enquiry from public sources, score it against what a good client actually looks like for you, and draft a first response. Sales time goes to the enquiries worth a call.

How we build

A Stockholm agent build, week by week

Two to four weeks from kickoff to handover. The order matters more than the tooling: measure first, prototype on real data second, close the loop third.

  1. 01

    Watch the workflow being done

    2 to 4 days

    We sit with the people who run the process today, and we measure it: how many cases, how long each takes, where they stall, and which exceptions actually recur. Most projects that fail do so because this step was skipped and the brief described the process as management believes it works rather than as it runs.

    You end up with: A measured baseline you can hold the finished system against, and a written list of the exceptions nobody had documented.

  2. 02

    Build the thin version on your real data

    3 to 5 days

    Not a demo on a curated sample. Your records, including the ones with missing fields and inconsistent formatting. This is where you discover that a third of the source rows lack something the workflow depends on, and it is much better to discover that in week one than in month three.

    You end up with: A narrow tool running on production-shaped data, and an honest assessment of whether the rest is worth building.

  3. 03

    Close the agent loop

    1 to 2 weeks

    Now the agent plans across steps, calls the tools it needs, and handles the cases the thin version could not. Human review gates go on every action that is expensive to undo. We build the evaluation set from your real cases at the same time, including the failures, because an agent with no evaluation set is an agent nobody can safely change later.

    You end up with: A working agent, an evaluation suite built from your own cases, and audit logging on every tool call.

  4. 04

    Hand it over properly

    2 to 4 days

    A runbook, a training session with the people who will operate it, and a documented path for what to do when it breaks. We do not make handover deliberately incomplete to keep you dependent on us. If you want us to keep operating it, that is a separate retainer you choose, not a trap you fall into.

    You end up with: Runbook, handover session, and the code and configuration in your own repository.

How we build

Positions we hold on every build

These are decisions we make the same way every time, because each one is a reason agent projects fail when it goes the other way.

Bounded autonomy by default

An agent starts read-only and draft-first. Actions that are expensive or awkward to reverse stay behind a human approval gate, permanently if that is the right answer. Full autonomy is something a system earns by demonstrating accuracy on your evaluation set, not a launch feature.

An evaluation set from your real cases

Built from your actual records, including the ones the agent gets wrong. Without it, nobody can safely change a prompt or swap a model six months later, which is how working systems quietly rot.

Every tool call logged

Each action, its inputs, and the decision behind it are recorded. This is what makes an incident investigable, and under FINMA, NIS2 or medical-device rules it is a requirement rather than a nicety.

Retrieval quality over model size

Most disappointing agents are not under-powered, they are under-informed. Getting the right context in front of the model reliably matters more than which model it is, and it is where the engineering effort usually belongs.

Your repository, your infrastructure

Code and configuration live in your repository and run on infrastructure you control. There is no VLIRTZ platform you have to keep paying for to keep your own workflow running.

One workflow before three

We decline company-wide assistant scopes. A single workflow, measured and shipped, tells you more about whether this approach works for you than any roadmap, and it is recoverable if the answer is no.

Pricing

What an AI agent costs in Stockholm

We quote in SEK for this market. Where you land depends on how many systems the agent touches, how usable your data already is, and how expensive a wrong action would be. Ask and you get a range on the first call, not the third.

Agent feasibility review

On request

We measure the workflow, assess whether your data supports it, and tell you whether an agent is the right answer. Includes a scoped build proposal.

Timeline: 1 to 2 weeks

Scoped agent build

On request

One workflow end to end, integrated with your tools, human review gates, evaluation set, audit logging, runbook and handover.

Timeline: 2 to 4 weeks

Additional workflow

On request

A second or third agent reusing the orchestration, retrieval and evaluation harness from the first.

Timeline: 2 to 3 weeks each

Sustain retainer

On request

Monitoring, drift checks, model and prompt updates, and a defined response time on failures.

Timeline: Rolling

Straight answers

What an AI agent will not do for you

Every one of these has ended a project somewhere. We would rather raise them before you sign than explain them in month two.

It will not fix a process nobody has agreed on
If two departments genuinely disagree about how a case should be handled, an agent forces that disagreement into the open rather than resolving it. That is useful, but it is a management outcome, not a technical one.
It will not rescue unusable source data
Retrieval over clean, structured records is straightforward. Retrieval over scanned documents, three competing sources of truth, and a field that has been wrong since a migration is where budgets disappear. Sometimes the honest recommendation is a data project first.
It will not be right every time
The question is never whether it makes mistakes, it is whether the mistakes are caught before they cost anything. That is what the review gates and the evaluation set are for, and it is why we measure the baseline first.
It will not maintain itself
Models change and your source systems change. An unmaintained agent degrades quietly rather than failing loudly, which is worse. Budget for maintenance or plan to retire it.

More about working with us in Stockholm

This page covers how we build agents. The Stockholm market page covers the rest: the regulators that shape a project there, the sectors we see most, and when we are the wrong partner.

AI Consulting in Stockholm

FAQ

AI agent development in Stockholm: common questions

How long does it take to build an AI agent?
A scoped single-workflow agent is typically two to four weeks from kickoff to handover, with a feasibility review of one to two weeks ahead of it if the use case is not settled. Anything quoted at under a week is a demo, and anything open-ended is a warning sign.
How much does AI agent development cost in Stockholm?
A scoped build is the usual entry point. What moves the number is how many systems the agent touches, how usable your data already is, how expensive a wrong action would be, and whether your team or ours operates it afterwards. Our pricing page publishes the bands and the drivers.
Will the agent act on its own or ask permission?
It starts read-only and draft-first, and anything expensive to undo stays behind human approval. Autonomy is earned by demonstrating accuracy on your evaluation set over real volume, not enabled at launch. For many workflows draft-and-approve is the permanent right answer, and it still captures most of the time saving.
What do we actually own at the end?
The code and configuration in your own repository, running on infrastructure you control, plus the evaluation set, the audit logging and a runbook. There is no VLIRTZ platform you have to keep paying for to keep your workflow running.
Which frameworks and models do you use?
We are deliberately not tied to one. Most builds use an orchestration layer such as LangGraph with a hosted model, and a vector store for retrieval, but the choice follows your data residency and latency constraints rather than our preference. Retrieval quality matters more to the outcome than which model sits behind it.
Do you work on-site in Stockholm?
Yes, and it is the default rather than an exception. We are based in Stockholms lan, so the observation sessions, kickoff and handover training happen in your office. That is where the undocumented exceptions surface, and they rarely surface over video.
Can you build agents that work in Swedish?
Yes. Handling Swedish-language input and output is straightforward, and workshops can run in Swedish. Delivery documentation is in English by default because it tends to survive team changes, but we will write it in Swedish if you prefer.
What if the agent gets something wrong?
It will, and the design question is whether mistakes are caught before they cost anything. That is what the review gates, the evaluation set built from your real failures, and the audit log on every tool call are for. We measure the baseline first so you can tell whether the finished system is actually better than the process it replaced.
Do we need an AI strategy first?
No, and most Stockholm clients do not have one. A specific workflow that costs you time or revenue is a better starting point than a strategy document, because it gives the first build something measurable to aim at.

This page was last reviewed on .

By market

Agent development in other markets

Each market page covers the workflows we are actually asked to build there, the regulators that shape the design, and pricing in the local currency.