Choosing an enterprise AI platform is a choice between three categories, not a long list of products: a cloud vendor's managed AI stack, a build on open-source frameworks, or a deployable platform you run yourself. The criteria that separate them are deployment posture, identity and access enforcement, audit depth, model substitutability and who operates it — not feature counts.

For institutions that cannot send data to a foreign cloud, the first category is eliminated before the evaluation starts, which changes the whole exercise.

What are you actually choosing between?

A cloud vendor's AI platform. Managed inference, retrieval, agent orchestration and observability from a hyperscaler or model provider. Fastest to start, deepest feature set, least operational burden. The constraint is where processing happens: sending a prompt containing personal data to a foreign service is a cross-border transfer, as set out in data residency and personal data law in Azerbaijan. For a supervised bank or a state institution here, that is usually decisive rather than a factor to weigh.

A build on open-source frameworks. Assembling orchestration, retrieval, tool integration and observability from components. Maximum control and no licensing constraint. The honest cost is that the components below are yours to build and maintain: identity propagation, policy gates, audit trail, evaluation harness, admin interfaces. Teams consistently underestimate this, because a working prototype is a weekend and a governable production system is a year.

A deployable enterprise platform. A product you run inside your own perimeter, with orchestration, policy and audit as built-in capabilities. The middle path: less flexible than a build, deployable where a cloud platform cannot go.

A fourth option that is not a platform: a point solution — an AI feature inside an existing application. Frequently the right answer for one use case and a dead end as a strategy, because each vendor's feature has its own model, its own data path and its own audit posture, and the institution ends up with six ungoverned AI systems instead of one governed one.

Which criteria actually separate platforms?

Ten, in the order they eliminate options rather than the order vendors present them.

1. Deployment posture. Can it run entirely inside your perimeter, including air-gapped? This is binary and it is first, because it eliminates whole categories. Ask specifically: which components require outbound connectivity at runtime, and what happens to updates in an air-gapped environment.

2. Identity propagation. Does the user's identity reach every retrieval and every tool call, or does the platform hold a service account and filter afterwards? The second pattern means any successful prompt injection inherits broad access. This is the most important security criterion and the one most often answered vaguely — the mechanics are in building secure AI workflows.

3. Access control integration. Does it enforce your existing entitlements, or does it maintain a parallel permission model? A parallel model will drift, and after it drifts the AI system's access is nobody's idea of correct.

4. Model substitutability. Can you change models — including to a locally-hosted open-weights model — as configuration rather than as a rewrite? Given how fast the landscape moves, an architecture that locks the model has an expiry date.

5. Audit depth. For a given output, can the platform show the requesting identity, the input, the retrieved context, the model and version, the tool calls with parameters, the policy evaluations and the approvals? Test this by asking for a real audit record, not a description. The requirement is set out in what the ISO/IEC 42001 audit asks for.

6. Policy and approval gates. Are they platform capabilities or application code? Approval workflows written per application diverge, and one of them will eventually be the one that skips the gate.

7. Tool and integration model. Are tools defined once and reusable, ideally through a standard interface, or rewritten per use case? The Model Context Protocol has become the common approach and its practical value is exactly this reusability.

8. Evaluation tooling. Can you maintain a question set, measure retrieval separately from generation, and re-run it on every change? Platforms without this force you to build it, and without it quality discussions stay anecdotal.

9. Multilingual behaviour. Retrieval and generation quality in Azerbaijani, not only English. This is measurable and it is frequently the criterion that separates otherwise similar options in this market.

10. Operating burden and who carries it. Who patches it, who monitors it, who is called at 2am. For on-premise deployments this is a staffing question that belongs in the evaluation rather than after it.

What should you not evaluate on?

Three criteria that consume evaluation time and predict nothing.

Model benchmark scores. Public benchmarks measure general capability. Your outcome is determined by retrieval quality on your corpus, and the correlation between benchmark rank and enterprise answer quality is weak — the reasoning is in why context matters in enterprise AI.

Feature count. Every platform in this category has a long list. The features that matter are the ten above, and most of them are architectural properties rather than checkboxes.

Demo quality. Demos run on curated corpora with cooperative questions. A demo tells you the platform can do the easy case, which was never in doubt.

How should the evaluation actually run?

A four-week structure that predicts production far better than a scoring matrix.

Week 1 — eliminate on deployment posture and identity. Two questions, asked precisely, in writing. Options that cannot run in your perimeter or cannot propagate identity are out, and this usually reduces a list of eight to a list of two or three.

Week 2 — build the evaluation set. A hundred real questions from real users, with the correct source document identified, in all the languages your organisation works in. This is reusable across every candidate and it is the most valuable artefact the evaluation produces — it outlives the decision.

Week 3 — run each remaining candidate on your own corpus. Not a vendor demo dataset. Measure retrieval precision separately from answer quality, per language. Expect the results to be worse than the demo suggested; that gap is information about how the platform behaves on messy real content.

Week 4 — test the boundary, not the model. Attempt prompt injection through a document. Attempt to retrieve something the test user should not see. Trigger an action that should require approval. Then ask for the audit record of everything you just did. A platform that cannot produce that record has failed the criterion that matters most in a supervised institution.

The output is not a score. It is a short document saying which candidates survived weeks 1 and 4, and how they compared on your corpus in week 3.

Which criteria are deal-breakers in this market?

Three, and they are worth stating plainly because they reorder the shortlist.

Inference must run inside the perimeter for supervised institutions handling personal data. Not because on-premise is inherently better, but because the transfer rules make the alternative hard to defend, as covered in sovereign AI. This eliminates most managed cloud platforms at week 1.

Logs and telemetry must stay inside too. A platform that keeps inference local and ships observability data abroad has relocated the problem. Ask explicitly where logs go; it is frequently a separate answer from where inference runs.

Azerbaijani-language performance has to be measured, not assumed. Azerbaijani is substantially less represented in general-purpose models and embeddings than English or Russian. A platform that performs well in English may perform noticeably worse here, and no vendor will volunteer this.

Build, buy, or somewhere in between?

The honest framing, since this is the decision behind the decision.

Build if you have a platform engineering team that will still exist in three years, a use case sufficiently unusual that no product fits, and an appetite for building identity propagation, policy gates and audit trails yourself. These are not exotic, and they are considerably more work than they look.

Buy a managed cloud platform if residency permits it and the operational burden is what you are trying to avoid. Where this option is available it is usually the fastest route to value.

Deploy a platform inside your perimeter if residency rules out the cloud and you would rather not build the control plane. This is where most supervised institutions in Azerbaijan land, and it is a constraint-driven answer rather than an ideological one.

There is no universally correct choice. There is a correct choice given your residency constraints, your engineering capacity and your audit obligations — and those three, not the feature comparison, determine it.

Disclosure

Yukon Labs builds and deploys HAVAA, which sits in the third category: an orchestration platform that runs entirely inside the customer's perimeter, including air-gapped, with identity propagation, policy gates and audit trail as platform capabilities.

That is a commercial interest and this article is written with it disclosed rather than concealed. The evaluation framework above is the one Yukon Labs uses in engagements, including where the conclusion is that a customer should build or should use a cloud platform — which is the right answer whenever residency permits it and the operational burden is what the organisation is trying to avoid.

The criteria are worth applying regardless of what you conclude, and weeks 2 and 4 are worth running even if the decision is already made, because the evaluation set and the boundary tests are what you will need in production anyway.

Key points

  • The choice is between three categories: managed cloud platform, open-source build, or a platform deployed inside your perimeter.
  • Deployment posture and identity propagation are the first two criteria because they eliminate categories rather than scoring them.
  • Test model substitutability, audit depth, policy gates as platform capabilities, tool reusability and evaluation tooling.
  • Do not evaluate on benchmark scores, feature counts or demos. None of them predicts production behaviour.
  • Run a four-week evaluation: eliminate on posture and identity, build a hundred-question evaluation set, run candidates on your own corpus per language, then test the enforcement boundary and ask for the audit record.
  • In this market, in-perimeter inference, in-perimeter logging and measured Azerbaijani-language performance are deal-breakers.
  • Build if you will still have the team in three years and are willing to build identity, policy and audit yourself; that work is larger than it looks.

Yukon Labs runs this evaluation as part of a readiness assessment, and deploys HAVAA on-premise where that is the conclusion. For the components any candidate must provide, see AI orchestration architecture; for the failures that follow a bad selection, common AI implementation mistakes.