Oridex

Service

Without the right
data foundation,
every AI
transformation fail.

Oridex is that foundation. It turns your documents – including the unstructured, scanned, and mixed-quality ones into a governed, retrievable layer inside your own infrastructure, on-premises or private cloud. Models are served inside that boundary, not called from a third-party API; a review step sits before anything goes live; and every result carries its source. Build the applications on top of it. Or let us.

oridex — data pipeline run
Sample data

Corporate archive · batch 04

1,284 files

File

Pages Status

msa_supplier_2019_signed.pdf

48 pages
Extracted

board_minutes_2014_scan.tif

12 pages
Extracted

annual_report_2025.docx

96 pages
Extracted

scan_20160812_lowres.jpg

1 pages
Retry file

supplier_register_q2.xlsx

22 pages
Extracted

1,281 extracted

3 need attention

Run report ready for review

6 metadata fields tagged

Status tracked per file, not per batch. The run report is reviewed before output is selected for retrieval. Interface shown with sample data.

ISO 27001 · CMMI Level 3 · PDPA-ready · On-premise / private cloud · Backed by DEHA Group (Singapore · Vietnam · Japan)
/ Why they fail

The model was fine.
The foundation wasn't there.

When AI projects stall, it is rarely the model – the LLM itself. It is the data layer underneath, and usually in one of four places.

01

The documents were never machine-readable.

Years of scans and files with no consistent structure. The pilot ran on a handful of documents someone cleaned by hand; production had a document set several orders of magnitude larger that nobody had touched – so it never left the pilot.

02

Nothing traced back to a source.

A result with no source document behind it cannot be checked – so nobody could verify it, and the business stopped relying on it.

03

Nobody reviewed the output before it went live.

Extraction quality was assumed rather than checked. The first close look came when a user reported a wrong answer – and trust did not recover.

04

The data had to leave to be processed.

The architecture assumed documents could be pushed to an outside service for processing. For a regulated organisation, that assumption ended the project at the security review.

That layer is where Oridex sits.

/ Where it sits

The layer between
your documents and your applications.

Oridex is not a chatbot. It is the platform a chatbot – or any other AI application – stands on. Build on top of it with your own team, with ours, or with Oriene AI, our answer application.

Applications

Oriene AI - answers with the source cited, served on top of Oridex
Multi-step AI workflows and agents
Downstream applications
Open API — any application can query it
Query the retrieval layer – ask a question, get the exact passage and the document it came from.

Oridex

Repository
Data pipeline
RAG pipeline
Model serving
Ingest, extract, review, index, serve – inside your infrastructure, models included.

Your environment

Documents
Document archives
Object storage
Infrastructure
Kubernetes cluster - on-premises or private cloud
Identity
Your identity provider
Nothing is replaced. Oridex is deployed onto your infrastructure and sits alongside your DMS and core systems.
/ What's inside

Four modules, one platform.

oridex - console
Product UI
The console as delivered. The four entry points on screen are the four modules below.

01 · Repository

Documents, governed

Buckets, files, document metadata and the templates that define what gets extracted.

02 · Data pipeline

Reads what you hold

Reads your PDFs, scans, images, Word and Excel files – including messy, mixed-quality ones – and turns them into clean, structured text, with the operational controls a large document set needs.

03 · RAG pipeline

Made retrievable

Turns reviewed content into AI-ready, retrievable data – chunked with sliding windows, embedded and indexed, validated before it goes live.

04 · Model serving

Models served inside your infrastructure

Your platform admin registers and serves a model once; pipeline owners simply select it. Open-source embedding and language models run in your cluster, not behind a third-party API.

/ Control

Two pipelines, each with a review step.

Both pipelines include a review step: the pipeline owner checks the run report and output before selecting it for the next stage. This is what lets you answer a compliance function without promising an accuracy number.

Data pipeline
documents → structured output

01

Upload documents

Into a pipeline with a bucket and a metadata template

02

Select model

The extraction model that suits the document type

03

Configure

Split size, page limits, concurrency, fields to extract

04

Processed output

Text, structured output, tagged fields, per-file status

05

Review

The pipeline owner checks the run report and output before selecting it for retrieval
RAG pipeline
structured output → retrievable

01

Select datasets

Completed data-pipeline output, checked by the pipeline owner

02

Select model

Embedding model selected for your language and document set

03

Configure

Chunking, vector collection, index settings

04

Processed output

Vectors written, dimensions validated

05

Review

Retrieval quality checked against your questions before go-live
 
/ Where it runs

Inside your boundary
and we tell you where it isn't.

01

Documents

Uploaded, or synced from your object storage.

02

Extract

Split, OCR, NLP, fields.

03

Review

Your team checks output before it moves on.

04

Index

Chunk, embed, write vectors and index.

05

Serve

Retrieval endpoints for your applications.

Models

Embedding and language models served from your own cluster.

Identity

Ships with Keycloak, federated to your identity provider.

Storage

Documents, extraction output and vectors on your infrastructure.

Deployment

Your Kubernetes cluster — your AWS or Azure region, or your own data centre.

Two things cross the boundary, stated plainly. Document extraction runs through an OCR provider configured for your deployment – a pluggable setting, changed without code – over egress you approve. And the first download of open-source models is a one-time pull, which can come from your own registry instead. Everything else – storage, embeddings, vectors, model serving and the applications querying them – stays inside your network. If your documents cannot leave, we’ll bring extraction in-cluster as scoped work on the platform, as it was for the deployment described below.

Your network

/ Who builds and runs it

Built by DEHA Group. Operated by you - by design.

DEHA GLOBAL is the Singapore entity of DEHA Group. Oridex is built, deployed and supported by the group’s shared engineering centre — the same organisation that has shipped production systems across Vietnam, Japan and Singapore for a decade.

200+

engineers in the shared delivery centre

600+

projects delivered since 2016

ISO 27001

information security, certified

CMMI L3

development and services, appraised

After handover

What you hold after handover - so you can operate it without us

Deployment manifests and Helm values – the infrastructure as code, in your repository
Automated test suite – API, interface, database and performance checks, so you can verify any change without us
Automated test suite – API, interface, database and performance checks, so you can verify any change without us
Smoke-tested staging and production – handed over working, not described

This is a design choice, not a limitation. A data foundation you cannot operate yourself is not a foundation. We stay on a support arrangement you scope; the delivery centre remains behind it.

oridex - dashboard
Product UI
What your operations team sees after handover: GPU capacity and utilisation, models being served, and every pipeline run by status.
/ In production
Delivered
Korea
Insurance

A decade of scanned records, made retrievable inside the customer's own data centre.

The situation

An insurance group holding more than ten years of policy, claims and legal records, largely scanned and in Korean. Public cloud tools had been ruled out: the files could not leave, and the volume did not fit.

What was built

  • Built on Oridex – the same four modules described above
  • Models and processing deployed inside the group’s own data centre
  • Recognition models selected for Korean rather than English-first defaults
  • Two applications on the same foundation – one for staff, one customer-facing
How the platform extends. The core is standard. The fully offline configuration, the Korean-language extraction models and the knowledge-graph layer were engineered on top of it for that deployment — which is exactly what a foundation is for. If your documents need the same, it is scoped as work on the platform, not assumed to ship with it.
/ How an engagement runs

Scope, evaluate, decide - then deploy.

One use case, one document set, one measure of success. We scope before we quote, and we deploy once the evaluation clears your criteria and your infrastructure prerequisites are in place.

01
Week 1

Scope

The documents, the questions that matter, the constraints. One use case, written down, with success criteria agreed before any build.

02
Week 2-3

Evaluate

Oridex runs on a slice of your own documents – on your cluster once prerequisites are in place, or on a sample you approve in an isolated staging environment agreed during scoping. Extraction and retrieval quality measured against your criteria, with your engineers alongside ours.

03
Week 4

Decide

You see the results on your own corpus. If it clears, we move to deployment with a scoped plan and a fixed prerequisite checklist.

04
Deploy

Install and hand over

Installed in your environment, smoke-tested, with manifests, run book and test suite handed to your team. Timing depends on your DNS, certificates and cluster access.

How it is priced

Two lines from the first contract: a platform licence, and the integration work to put it into your environment. You see what recurs and what does not.

Evaluation

Priced as a fraction of the first production year, and credited against it if you proceed. Scoped to one use case so the number stays small.

What we need from you

A Kubernetes environment (Rancher-managed or compatible), S3-compatible object storage, and a named owner on your side. GPU capacity is sized during scoping, not assumed. We share the full prerequisite checklist in week one.

/ Talk to us

Bring the problem.
Leave with an architecture.

Forty-five minutes with Brian Dang, CEO of DEHA GLOBAL. You describe where your AI initiative is stuck. We map it against Oridex and tell you – plainly – whether it fits, what it would take, and what it would not solve.

Brian Dang

Chief Executive Officer, DEHA GLOBAL · Singapore

Brian Dang leads DEHA GLOBAL and works directly with CIOs and CTOs on document and data foundations for AI. When the conversation gets into cluster topology or model sizing, a solution architect from the delivery centre joins the call.

0–15 min

You describe the documents, the initiative, and where it stalled

15–35 min

We map it against Oridex — architecture, boundary, what leaves and what doesn’t

35–45 min

Straight answer: fit or not, and the next step if yes

Book a meeting with Brian Dang

Three questions, because they decide whether the meeting is worth your time. We reply within two working days with proposed slots. By submitting you agree to our Privacy Policy (Singapore PDPA).

Not ready for a meeting? Get the technical overview by email — architecture, deployment boundary, prerequisites.

/ FAQ

What CIOs ask first.

Where does our data go?

Your documents, extraction output, vectors, and served models stay inside your environment- on-premises or in your own cloud region. Document extraction runs through an OCR provider configured for your deployment, over egress you approve; which provider, and exactly what leaves in that call, is agreed during scoping and named in the Data Processing Agreement. If your documents cannot leave, we’ll bring extraction in-cluster as scoped work. A DPA, the sub-processor list, and support for your own penetration test are available under NDA.

No. Oridex is the retrieval layer underneath — structured, searchable knowledge and the endpoints to query it, with results carrying the source document and passage they came from. Build on it with your own team, or use Oriene AI, our answer application that runs on top.

A Kubernetes environment, object storage and the usual platform services. GPUs are optional and depend on whether you serve models locally and at what size. We size this with your infrastructure team during scoping and give you the prerequisite checklist up front.

You check it at a review step in both pipelines: the pipeline owner checks the run report and output before selecting it for retrieval. We measure extraction and retrieval on your own documents during evaluation, against criteria you set – rather than quoting an accuracy number that means nothing on your corpus.

Deployment and support are done by named engineers from the group’s delivery centre in Vietnam, through the remote-access path your security team approves, with sessions logged on your side. The access arrangement is agreed and documented before any work begins, so it can go through your outsourcing review up front rather than after.

Your team, with a documented handover: architecture, run book, deployment manifests, and an automated test suite. We stay on the support arrangement you scoped, and the group’s delivery centre supports it.

Scans are read with OCR rather than re-keyed, and the embedding model is selected for your language. Rather than claim a quality level, we measure it on your documents during evaluation – before anything is committed.

Give your AI a foundation
it can stand on.

Book a working session with Brian Dang, or get the technical overview – architecture, deployment boundary and prerequisites.