AI built around a task, not around the technology

AI development means building software that uses language models or machine learning to do a specific job: answering customer questions, routing enquiries, extracting data from documents, or automating a workflow. MBGDesk scopes this against a task that currently costs you time, then builds and measures against that task.

Most AI projects fail for a reason that has nothing to do with the model. They start from the technology and work backwards to a use case, which produces demonstrations that impress in a meeting and are quietly abandoned within a quarter. The projects that survive start from a task someone is doing repeatedly and expensively.

The specific risks are well known by now and routinely ignored. A model asked a question outside its knowledge will answer anyway, fluently and wrongly. A chatbot with no escalation path traps a frustrated customer. An automation with no audit trail makes decisions nobody can explain afterwards. And a system with no evaluation set cannot be improved, because nobody can tell whether a change made it better.

Who this is for

When this work makes sense

  • Businesses answering the same customer questions repeatedly across email, chat and phone
  • Teams manually extracting data from invoices, forms, applications or CVs
  • Companies with a support queue where triage takes longer than resolution
  • Organisations wanting to use their own documents as a knowledge base
  • Products that need AI as a feature rather than as a separate tool
How to choose

Questions worth asking any supplier

Can you name the task and what it currently costs in hours or headcount? If not, the project has no success condition.

What happens when the model is wrong? A system without a defined failure path will fail in public.

Where does your data go? Some providers train on submitted data by default. This matters for anything confidential.

Is there an evaluation set? Without test cases and expected answers, 'improving the prompt' is guesswork.

Is there a human escalation route? Every deployment that handles customers needs one, and it should be easy to reach.

What you get

Deliverables

A scoped use case with a stated success measure agreed before build

The working system: agent, chatbot, automation or integration

Retrieval over your own content where the task needs your knowledge rather than general knowledge

Guardrails: scope limits, refusal behaviour, and a defined escalation path to a person

An evaluation set of real cases with expected outcomes, so changes can be measured

Logging and an audit trail of what the system did and why

Documentation your team can operate from, including how to update the knowledge base

Built with

  • Large language model APIs
  • Retrieval-augmented generation
  • Vector databases
  • Workflow automation
  • Node.js
  • Python
  • Webhook and CRM integrations
How we work

Process

AI work runs an evaluation stage that other projects do not have. Before launch we assemble a set of real inputs with known-good outputs and score the system against them. That set becomes the regression test for every subsequent change. Without it, tuning a prompt is indistinguishable from guessing, and nobody can tell whether last week's change helped.

The standard nine stages — discovery, strategy, UX/UI, development, content and SEO, testing, launch, measurement, continuous improvement — apply to every project. The paragraph above is what differs for this one.

Pricing

What moves the number

Pricing depends on scope, functionality, integrations and content requirements. These are the factors that change it most:

  • How well-defined the task is — vague scope is the main cost driver
  • Whether the system needs your own documents, and what state that content is in
  • Integration surface: CRM, helpdesk, WhatsApp, internal systems
  • Accuracy requirement, which drives how much evaluation and iteration is needed
  • Ongoing model API costs, which are yours and scale with usage
  • Whether the system needs monitoring and retraining after launch
Comparison

Off-the-shelf tool, custom build, or neither

Off-the-shelf AI tool

Fits: Common tasks with standard workflows — meeting notes, first-line chat, content drafting

Limits: Limited control over accuracy, data handling and tone; you fit your process to the tool

Custom build

Fits: Tasks specific to your business, your documents or your systems

Limits: Higher cost and needs your subject-matter time to define what correct looks like

Not using AI

Fits: Tasks that are rare, high-stakes, or where a wrong answer is expensive

Limits: Nothing — this is often the right answer and we will say so

Proof

Case studies

No published AI case study yet. The field is full of demonstrations presented as deployments; we would rather publish nothing than add to that.

MBGDesk operates from 4 offices — India, United States, United Kingdom, United Arab Emirates. We hold no industry certifications and do not imply otherwise.

Questions

Frequently asked questions

Will AI replace our support team?

In our experience it changes what they spend time on rather than removing them. Systems that handle routine questions well still need people for the cases that matter, and a deployment with no human escalation path tends to produce angrier customers rather than fewer.

What if the AI gives a customer wrong information?

This is the central design question, not an edge case. We scope what the system is allowed to answer, build in refusal for anything outside that, ground answers in your own content rather than the model's general knowledge, and make escalation to a person easy. We also log everything so a wrong answer can be found and corrected.

Is our data used to train the model?

It depends entirely on the provider and the tier. Some train on submitted data by default. We establish this before choosing a provider and tell you what the terms actually say, which matters if you handle anything confidential.

How accurate is it?

We cannot answer that in the abstract, and any provider who quotes an accuracy figure before seeing your task is making it up. We build an evaluation set from your real cases and measure against it, so the answer is specific to your use.

What does it cost to run?

Model API costs are usage-based and belong to you, in your own provider account. We size them against expected volume during scoping so the running cost is understood before you commit to a build.

Can it work with WhatsApp?

Yes, through the WhatsApp Business Platform. This is a common request in India in particular, where WhatsApp is frequently the primary customer channel rather than an afterthought.

Do we need AI at all?

Often not. If a task happens rarely, or a wrong answer is expensive, or a simple rule would do the job, we will tell you that. Talking clients out of AI projects has been a recurring part of these conversations.

Talk to us about this

Tell us what you are trying to build and we will come back with scope, approach and an estimate.