IT security intelligence. Since 2006.Cloudflare services ↗
AI, automation & data

Custom AI models, AI-as-a-Service & MLOps

Develop and operate AI systems around your data, tasks, and constraints. Evaluate whether retrieval, adaptation, or model training is the right approach before investing in complexity.

Discuss your requirements
Researcher evaluating a machine-learning model and document retrieval results

Select the right approach for the task

A custom system does not always need a custom-trained model. We compare existing models, retrieval-based systems, fine-tuning, and specialist machine-learning approaches against representative tasks. Evaluation considers quality, language coverage, data handling, latency, and cost.

Work can include data preparation and labelling, synthetic data where appropriate, model adaptation, and API integration. The rights to use training material and the handling of personal or confidential information are assessed before development.

Plan for operation from the beginning

Models and datasets change. A production plan needs versioning, evaluation datasets, deployment procedures, observability, and a way to roll back. MLOps processes make these changes reviewable and help the team understand why output quality has shifted.

Hosted AI services and AI-as-a-Service can be scoped around required availability, access, usage limits, and support. Model updates and retraining are controlled processes, with review criteria and an identified owner rather than an assumption that more data always improves results.

Prepare data that represents the actual task

We assess provenance, permitted use, quality and coverage before adapting a model. Training, development and evaluation material are separated to reduce misleading results. Examples should represent difficult cases as well as common ones, including Greek and English requirements where relevant. We document annotation guidance and disagreement, examine sensitive fields and establish a baseline that a customised approach must improve upon.

Compare approaches using an agreed evaluation

Candidate models are tested on the same task set with defined scoring and human review. Fine-tuning may help with specialised behaviour or output formats, while retrieval can provide changing source knowledge. Neither approach is assumed to be best in advance. The comparison considers accuracy, failure modes, response time, resource needs, provider dependencies and total operating cost. The chosen architecture includes an explanation of the trade-offs.

Engineer release, rollback and monitoring

The delivered system versions the model or provider selection, prompts, data transformations and evaluation suite. Release checks compare the candidate with the accepted baseline and include important failure cases. We define how to detect a degradation, investigate it and roll back a change. For hosted services, access boundaries, usage controls and availability requirements are included in the operating design. Retraining and upgrades follow an explicit approval and validation process.

Make the system transferable to your team

Documentation covers interfaces, infrastructure, deployment steps, evaluation results and support responsibilities. We distinguish assets the client owns from provider-hosted components and third-party licences. The handover includes the skills and access needed for routine operation, plus an exit or migration view where dependence on a supplier matters. This gives technical and procurement teams a practical basis for maintaining the service.

Choose between prompting, retrieval and fine-tuning

A model that answers from changing company documents may need better retrieval, while a narrowly defined classification task may benefit from task-specific examples or model adaptation. We examine the failure that needs to be corrected before choosing a technique. Evaluation uses representative inputs and a comparison with a simpler baseline, so additional complexity must demonstrate a useful improvement.

Test Greek and English with real task variation

A bilingual application should be evaluated with the terms, abbreviations, document formats and mixed-language inputs its users encounter. We assemble an agreed test set and separate training or development examples from the material used to assess performance. Errors are reviewed by type, including missing information, incorrect classifications and outputs that look plausible but do not answer the task.

Select a hosting model with operational consequences in view

A managed model API, a dedicated deployment and self-hosted inference create different responsibilities for capacity, access, maintenance and cost. We assess the relevant data flows and provider terms with your team, then document the selected arrangement. The operating plan includes version changes, evaluation before release and what should happen if a service is unavailable or produces unacceptable results.

What you receive

  • Model and architecture evaluation
  • Dataset preparation and usage requirements
  • Versioned pilot, evaluation suite, and deployment documentation
  • Monitoring, change control, and operating plan
  • Method-selection comparison with representative bilingual evaluation cases
  • Model operating plan covering versions, fallback and release criteria

Common questions

Do we need to train a model from scratch?

Often not. An existing model with well-designed retrieval or a targeted adaptation can be a more practical starting point.

How will model quality be measured?

We agree on representative examples and acceptance criteria, including failure cases, language performance, review effort, and operational cost.

How much labelled data do we need?

There is no useful universal quantity. It depends on the task, label quality, variation and chosen method. A sample review helps determine whether the next step should be data preparation, evaluation or model adaptation.

Can you evaluate a model supplied by another vendor?

Yes. A defined evaluation can compare the vendor’s system against representative tasks, agreed acceptance criteria and a baseline. Access, confidentiality and permitted testing methods must be established first.

What’s your next
technology challenge?

Talk to our team