Applied artificial intelligence research
The question that reaches us is almost never "build me this tool" but "can a model do this, on our data, at a cost that holds up". Answering it is research, not development.
The evaluation set first
One to three hundred real cases, drawn from your own documents, with the expected answer written by hand. It is slow, it is thankless, and it is the only thing that makes the rest arguable. Without that set, any comparison between two approaches reduces to an impression and to how much you trust the person voicing it.
The set stays with you. It serves the first decision, then every later model change, and it lets you check a future supplier without starting over.
- Real cases from your documents, not fabricated examples
- Expected answers written with your own subject experts
- Hard cases and traps included on purpose
- The set is delivered and usable without us
Comparing approaches, not defending ours
We put plain prompting with context, retrieval augmented search, fine tuning an open model and calling a proprietary model in competition. For each we measure accuracy on the evaluation set, cost per call, latency and the rate of expensive mistakes.
The result fits in a table and a recommendation. Sometimes it concludes that no approach justifies the project, and we say so before any development has been committed.
- Four families of approaches put in competition
- Accuracy, cost per call, latency, expensive error rate
- Open models run on our own infrastructure, in Switzerland
- A written, measured report delivered whatever the conclusion
From verdict to system
When the conclusion is favourable, production reuses the evaluation set as a regression test. Every change of model, supplier or prompt is measured before it ships. That is what prevents silent drift, the moment a system keeps running while quietly producing worse results.
- The evaluation set becomes a regression test
- Measurement before every model or supplier change
- Drift tracked over time
- Reversibility assumed: going back stays possible
Applied artificial intelligence research
How long does a feasibility study take?
Three to six weeks in most cases, a good half of it spent on the evaluation set. The deliverable is a written report with the numbers, the limits observed and a recommendation, including when it is negative.
Do we have to hand over all our data?
No. A study runs on a representative sample, and that sample can be pseudonymised. Processing can run entirely on our Swiss infrastructure if your obligations require it.
What if the verdict is negative?
You keep the report, the evaluation set and the written argument that lets you close the subject internally. It is less exciting than a project launched, and far cheaper than a project abandoned after a year.
AI applications
Language models put to work on bounded, measured and reversible tasks.
In the division AI applications
Let us talk about what you want to build.
Describe your situation in a few lines. If it falls outside what we do well, we will say so immediately.
Get in touch