- Manufacturers
- Security and facilities teams
- Media teams
- Field operations
Computer Vision, Speech & Edge AI
Vision and speech systems for inspection, recognition, transcription, monitoring and edge-based inference in operational environments.
- Automated visual quality checks
- Searchable speech and media
- Faster incident detection
- Low-latency inference near equipment or users
- Image, video and audio data preparation
- Detection, classification and segmentation
- Speech-to-text and voice analytics
- Edge device and model optimization
- Human review and privacy controls
- Performance and environmental testing
- Discovery and approved solution requirements
- Configured or developed product release
- Integration and data-migration outputs where included
- Functional, security, performance and user-acceptance evidence
- Administrator and user documentation with training
- Go-live, warranty and support-transition package
- Quality inspection
- Video analytics
- Speech transcription
- Document and label recognition
- Edge appliance deployment
- Edition, modules and user/location count
- Hosting, environments and availability target
- Data migration and integrations
- Configuration versus custom development
- Security, compliance and assurance scope
- Training, support and service level
Content on this page comes from the governed ARRIX catalogue record DAI-07; pricing is confirmed only through a reviewed quotation.
Deep Learning & Multimodal AI
Deep Learning & Multimodal AI designs, trains and deploys neural network systems for data that ordinary analytics cannot read - images, video, natural language, speech, sensor sequences and combinations of them - and is proposed only where the problem genuinely requires it.
The problem this answers
The organisation holds information its systems cannot use: photographs of assets and damage, scanned and free-text documents, recorded calls, footage from cameras it already owns, sensor traces that a threshold cannot interpret. People read it one item at a time, and the size of the backlog quietly decides what gets looked at and what does not. At the same time there is pressure to adopt deep learning, often without a stated problem, and the real risk is buying a heavy system to solve something a far simpler method would have handled.
Data the business already owns becomes usable input
Models read images, documents, audio and video directly, so archives and existing camera feeds stop being storage cost and start feeding a process.
Human attention moves to the cases that need judgement
The system handles routine perception across the whole volume; people handle the exceptions, the disputes and the appeals it routes to them.
The same criteria are applied to every item
A model does not tire, change shift or drift in strictness over an afternoon, so variation between inspectors or reviewers is removed from the routine cases.
Work that was impossible at volume becomes possible
Every unit can be inspected rather than a sample, and every call reviewed rather than a handful, because the marginal cost of one more item is compute rather than labour.
Cost is contained because the architecture is sized to the problem
Pre-trained backbones, smaller task-specific models, distillation and quantisation are proposed before large architectures, and a classical method is proposed where it is sufficient.
The system can run where the data already is
On-premises and edge deployment keep confidential images, recordings and documents inside the environment where bandwidth, latency or confidentiality require it.
Where this changes the outcome
Deep learning earns its cost in one specific situation: the signal lives in raw perception or in sequence, and no set of hand-written rules or tabular features captures it. In that situation nothing simpler will do, and the capability changes what the surrounding process is able to do at all. Outside it, a neural network is a slower, more expensive and less explainable route to an answer a smaller model would have produced, and it commits the organisation to specialist skills, hardware and retraining for years afterwards.
| Organisation | Need | What ARRIX does | Outcome |
|---|---|---|---|
| A manufacturer inspecting product visually on the line | Catching defects consistently at line speed, on every unit rather than a sample | Visual inspection models trained on the client's own defect images, deployed at the line | Every unit is examined against the same criteria, and inspectors adjudicate the flagged and borderline cases instead of watching everything. |
| An insurer handling claims supported by photographs | Triaging and pre-assessing image evidence before a handler opens the file | Image classification and damage detection feeding the claims workflow, with the handler deciding | Straightforward claims are routed and prepared automatically while complex ones reach a human with the evidence already summarised. |
| An organisation whose obligations sit inside contracts and scanned documents | Finding clauses, dates, parties and amounts across a body of documents nobody has time to read | Document understanding combining layout, vision and language models over the client's own corpus | Obligations and terms become searchable structured data, with every extraction linked back to the page it came from. |
| A contact centre with recordings nobody listens to | Understanding what customers are actually calling about, across all calls rather than a sampled few | Speech recognition with task-specific classification of intent, outcome and compliance events | Call reasons and compliance exceptions are measured across the whole volume, and quality review targets the calls that matter. |
| A utility or infrastructure operator surveying assets by drone or vehicle | Finding the defects in footage faster than engineers can watch it | Object detection and condition classification over survey imagery, with confidence-ranked review | Engineers review a ranked shortlist with locations attached rather than working through hours of footage. |
| An operator whose equipment fails in ways no threshold catches | Recognising a failure signature in the shape of a sensor trace over time | Sequence models over multivariate sensor history, alongside the existing rule-based alarms | Emerging faults are detected from pattern rather than from a single reading crossing a limit, with the alarm system kept as the floor. |
| A site with existing cameras and a safety or throughput question | Counting, detecting and alerting on activity without sending video off site | Edge-deployed detection models sized to on-site hardware, with only events leaving the site | The question is answered locally, and raw footage of people never leaves the premises. |
The path, stage by stage
- Images
- Video
- Documents
- Speech
- Text
- Sensor sequences
- ARRIX DEEP LEARNING SYSTEM
- Detections and classifications
- Extracted structured data
- Alerts and escalations
- Human review queue
Qualify
Define the task, and test whether a rule, classical model or ordinary analytics already answers it.
Feasibility finding, including the simpler-method result
Data
Assess the raw material: volume, variability, labelling state, consent and what more would cost.
Data and labelling assessment
Architect
Select the architecture against data type, latency, hardware, explainability and where it must run.
Architecture decision with trade-offs stated
Train
Fine-tune or train with tracked experiments, reproducible configuration and recorded data snapshots.
Trained candidate models
Evaluate
Test on conditions held out by site, device or period; analyse errors by class and input quality.
Evaluation report and model card
Optimise
Quantise, prune or distil and convert to the target runtime until it fits the hardware and the latency budget.
Deployment artefact with measured cost per item
Deploy
Release to cloud, on-premises or edge with human review and escalation where the consequence warrants it.
Running inference service with review workflow
Monitor
Watch input drift, confidence distribution and escalation rates, and retrain on the agreed trigger.
Monitored system with a retraining path
For whoever has to sign it off
Open only what you need. Nothing here is hidden from print or from a browser without JavaScript.
What ARRIX delivers
- A feasibility assessment that includes testing whether ordinary analytics or classical machine learning already solves the problem, and reporting the result either way
- Data assessment: how much labelled material exists, how variable it is, and what obtaining more would cost
- Architecture selection with the trade-offs stated, including model size, latency, hardware, explainability and operating cost
- Transfer learning, fine-tuning or training from scratch, chosen by what the available data justifies
- Training pipeline, experiment tracking and reproducible runs rather than one-off experiments
- Evaluation against held-out conditions with error analysis by class, site, device and input quality
- Optimisation for the target hardware: quantisation, pruning, distillation and runtime conversion
- Deployment to cloud, on-premises or edge, with throughput and cost per item measured before go-live
- Human review and escalation workflow wherever the consequence of an error is material
- Monitoring for input drift and performance decay, with a defined retraining path and owner
What you receive
- Feasibility finding, including the comparison against a simpler method and its recorded result
- Labelled dataset specification and the evaluation protocol agreed before training
- Trained models with recorded evaluation and error analysis by segment and condition
- Model card stating intended use, tested conditions, known failure modes and inputs that are out of scope
- Optimised deployment artefact built for the target runtime and hardware
- Inference service or edge deployment package, with measured throughput and cost per item
- Human review, override and escalation workflow where the decision requires one
- Monitoring, alerting and retraining runbook with handover documentation
Technical benefits
- Architecture selected against data type, latency budget and deployment location rather than against novelty
- Transfer learning and pre-trained backbones used wherever they reduce the volume of labelled data required
- Model size, quantisation and runtime chosen so the system fits the hardware it must actually run on
- Evaluation on data held out by site, source, device or period, so the reported result reflects generalisation rather than memorisation
- Failure modes characterised: what it gets wrong, on which inputs, how often relative to the alternative, and what happens next
- Defined behaviour for inputs outside the trained distribution, including abstention and escalation rather than a confident answer
- Serving path engineered for throughput and cost per item, measured before go-live rather than discovered after
- Reproducible training: data snapshot, configuration, seed and environment recorded so a model version can be rebuilt
Governance and security
- Consent, licensing, privacy basis and data residency confirmed for images, recordings and documents before training begins
- Faces, voices and personal identifiers minimised, blurred or removed wherever the task does not require them
- On-premises and edge deployment offered where raw material may not leave the site, so only events and results travel
- Evaluation carried out across the conditions and population groups the system will meet, with the per-segment results recorded rather than averaged into one number
- Human review, override and appeal retained wherever an output affects a person
- Training data snapshots, model versions and configurations recorded so any output can be reconstructed and defended
- Inference endpoints access-controlled, with inputs and outputs logged to the extent privacy obligations permit
- A written statement of the inputs and purposes the model is not to be used for, carried in the model card
Technology options
| Area | Candidates | How it is chosen |
|---|---|---|
| Vision | Convolutional architectures, Vision transformers, Object detection and segmentation models, Pre-trained backbones with a task-specific head | A pre-trained backbone fine-tuned on the client's own images is the usual starting point, because it needs far less labelled data than training from scratch. Naming a technology here states what may be proposed; it is not a vendor relationship. |
| Language and sequence | Transformer encoders for classification and extraction, Sequence models for time-ordered sensor data, Speech recognition and speaker separation models, Small task-specific models in preference to a general large model where the task is narrow | This offering covers task-specific models that classify, extract or detect. Open-ended generation and assistants belong to the Generative AI offerings, and the two are frequently combined. |
| Multimodal | Joint image and text models, Fusion of separate per-modality encoders, Staged pipelines where one modality gates the next | A staged pipeline is easier to debug and cheaper to run than a single joint model, and is proposed first unless the signal genuinely lies in the interaction between modalities. |
| Frameworks and training | PyTorch, TensorFlow, Experiment tracking and a model registry, Distributed training where dataset size requires it | Chosen against what the client's team can maintain after handover, and against any existing standard they already run. |
| Deployment and optimisation | Portable runtimes such as ONNX Runtime, Quantisation and pruning, Distillation into a smaller model, GPU inference servers with batching, Embedded and edge accelerator modules | The deployment location is an architecture input, not a later step: a model that cannot fit the available device is a design failure rather than an optimisation problem. |
When you may not need this
Most business questions do not need deep learning, and ARRIX does not sell it where ordinary analytics or classical machine learning would answer the same question more economically. If the data is tabular - transactions, ledger entries, service records, sensor readings already reduced to features - then a regularised statistical model or gradient-boosted trees will usually match a neural network and will be cheaper to build, faster to run, easier to explain and simpler for your team to operate afterwards. If the task can be written down as a rule someone in the business already applies, write the rule. Deep learning becomes the right answer when the signal lives in raw perception or in sequence - the appearance of a component, the wording of a clause, the sound of a fault, the shape of a trace over time - and no set of hand-written features captures it. Where a client asks for deep learning and the assessment shows a simpler method is sufficient, ARRIX records that finding and proposes the simpler method, even though it is a smaller engagement.
Answered plainly
Our board has asked for deep learning. Will you just build it?
Only if the problem needs it, and establishing that is the first part of the engagement rather than an afterthought. A simpler method is tested against the same task and the result is recorded, so the decision is evidence rather than opinion. If ordinary analytics or classical machine learning answers the question, ARRIX will say so and propose that instead, even though it is a smaller piece of work, because deep learning carries costs that continue long after the build: specialist skills, hardware, retraining and a system that is harder to defend when someone challenges an output. Where the assessment shows the signal really is in images, language, audio or sequence, the deep-learning system is built and the reason it was necessary is written down.
How much labelled data do we need?
There is no universal number, and anyone offering one before seeing the task is guessing. It depends on how distinct the classes are, how rare the important case is, how much variation your environment produces, and whether a pre-trained backbone applies to your data type, which for common image, text and audio tasks reduces the requirement considerably. The practical method is to label a small representative pilot, measure whether performance still improves as the set grows, and stop when it stops improving. The labelling programme itself, including guidelines and quality checks, is covered by Training Data & Feature Engineering.
Can it run without sending our images or recordings to a cloud service?
Yes, and where confidentiality, residency or bandwidth require it that becomes an architecture input rather than a preference. Running on your own servers or on a device at the site constrains model size and latency, which is why the deployment location is decided before the architecture rather than after. Quantisation, pruning and distillation are used to fit the model to the hardware available, and the trade-off against accuracy is measured and shown to you rather than assumed. In an edge deployment only events and results leave the site, and the raw footage or audio stays where it was captured.
How accurate is it?
Accuracy is a property of a specific model on a specific dataset under specific conditions, not a property of a technique, so a number quoted before your data has been seen would be invented. What is agreed in advance is the evaluation protocol: which data is held out and by what criterion, which errors are counted, and how the segments are broken down. Results are reported per class and per condition rather than as one headline figure, because a single average hides the failure that matters. The decision threshold is yours to set, because it trades missed cases against false alarms and only the business can price that.
What happens when the model sees something it has never seen before?
That behaviour is designed rather than left to chance, because a model that answers confidently on an unfamiliar input is a design failure. Confidence thresholds and out-of-distribution checks route uncertain inputs to a human queue instead of forcing an answer. Escalation rates are monitored, and a rise in them is treated as a signal that conditions have changed and retraining may be due. Where the consequence of an error is material, human review stays in the loop by design and is removed only once the failure modes have been characterised, if at all.
Is this the same as the generative AI you offer?
No, although they use related architectures and are often combined in one system. This offering builds task-specific models that recognise, extract, classify or detect, and their output is a structured result you can act on and measure. The generative offerings build applications that produce language grounded in your knowledge, and they are evaluated in a different way. A typical combined system uses a model from this family to read a document or a recording, and a generative one to summarise or answer questions over what was extracted.
What ARRIX needs to know
These are the questions a scoped proposal answers. Bring what you can; the rest is established in the conversation.
- What exactly must the system recognise, read or predict, and who acts on the output?
- Has anyone tested whether a rule, a classical model or ordinary analytics already answers it?
- How much labelled material exists today, and who in your organisation is qualified to judge a correct label?
- How variable are the inputs: lighting, camera position, document layout, accent, language, sensor placement?
- Where must inference run - cloud, your own servers, or on the device - and what forces that choice?
- What is the cost of a wrong answer, and who reviews the cases where it matters?
- Are the images, recordings or documents subject to consent, privacy or data residency constraints?
- What hardware is available, and is there budget to operate and refresh it after the project ends?
- How quickly must an answer come back, and what happens to the process while it waits?