UnitedEdge AI Technical Blog Banner
UnitedEdge® AI is a 3U appliance built as two mirrored halves, so any single component can fail without taking down inference, VMs or networking. This guide is for infrastructure architects, platform engineers and security teams evaluating it for branches, plants, clinics and other sites without local IT.
The design goals are:
  • Sovereign by construction. Models, prompts, outputs and embeddings never leave the appliance.
  • No single point of failure. Two of everything that matters, running active-active.
  • Lights-out operation. Zero-touch enrollment, fleet policy and remote remediation, with no one on site.
  • One platform. Inference, VMs, Kubernetes, serverless and software-defined networking in the same box.

Hardware: two halves, one appliance

Each side of the chassis carries an AI Engine, a Compute Engine, its own network handoffs and its own power feed. The two sides run active-active and back each other up.
hardware
UnitedEdge® AI hardware layout · 2 sides, 4 component pairs
The AI Engines serve models; the Compute Engines run everything else. Storage is clustered NVMe shared across both sides. The chassis is ruggedized to ship: connectors are locked and heavy components are shock-mounted.

Specifications

Item Specification
Form factor 3U, 19-inch rack, 23 in deep
AI memory 1 TB total (512 GB per AI Engine)
Compute 64 cores across two Compute Engines
System memory 1 TB ECC RAM
Storage 2 × 3.84 TB clustered NVMe
Network 4 × 25 GbE for LAN and WAN (two handoffs per side), LTE backup for management
Power 2 independent feeds, about 1.4 kW peak
Out-of-band Remote console and remote power control
Physical security Lid, bezel and shock sensors

Inference: models, memory and endpoints

Model catalog. Open-weight models are curated, security-scanned and license-reviewed by UnitedLayer, covering frontier LLMs, multimodal and vision, coding, speech, embeddings, industry-specific and small models from US, EU and Chinese publishers. Tenant policy can allow or block models by origin, license or workspace, and enforces it at deploy time. You can also upload your own fine-tunes. Licenses vary by model.
Sizing. Approximate footprints from the catalog show how the 1 TB of AI memory fills:
Model Type Approx. memory Placement
DeepSeek V4 Pro Frontier, 1.6T MoE ~880 GB Spans both AI Engines
GLM-5.2 Frontier, 744B MoE ~420 GB One AI Engine
Llama 4 Maverick Multimodal, 400B MoE ~225 GB One AI Engine
DeepSeek V4 Flash Frontier, 284B MoE ~160 GB One AI Engine
gpt-oss-120b Frontier, 117B MoE ~65 GB One AI Engine
Gemma 4 31B Vision, 31B ~18 GB One AI Engine
Whisper large-v3 Speech, 1.5B ~3 GB One AI Engine
In practice that means two frontier models plus many small ones, or one very large model across both engines.
Endpoints. Each deployed model is served as a private, OpenAI-compatible endpoint behind one gateway and one API key per tenant. Existing code changes one URL:
Failover path. Each endpoint has a defined path: local AI Engine, then the peer AI Engine, then optionally the UPC AI Factory in your core. Traffic auto-routes to the peer engine above 85% load. Where the data boundary is set to sovereign, prompts never leave the site.

Cloud native and networking at the edge

The Compute Engines run four workload types from IT-approved blueprints, with placement, HA and policy applied automatically:
Blueprint What it runs Availability
Virtual machine Linux or Windows from approved images Live migration and HA restart across Compute Engines
Container app Managed Kubernetes, Helm charts, private registry Pods spread across both Compute Engines
Serverless function Event, schedule or API triggers; scales to zero Runs on either side
AI model endpoint Open-weight models, OpenAI-compatible API Peer AI Engine and AI Factory failover
Before a launch, policy checks confirm placement fits site capacity, the HA reserve on the peer node is preserved, network policy is met and an approver is recorded (auto-approval where policy allows).
Networking is software-defined: virtual firewalls, load balancers and routers, configurable per tenant and highly available across both sides. Business units get isolated compute, storage, network and AI.

Failure modes

What fails What happens Service impact
Power feed Appliance runs on the other feed; a facilities ticket is opened None
AI Engine Endpoints re-route to the peer AI Engine, overflow to the core AI Factory if configured Higher latency at peak; no failed requests by design
Compute Engine VMs, pods and functions restart on the surviving Compute Engine Brief restart for VMs; pods and functions reschedule
Network link Traffic moves to the other side's handoff; firewall and load balancer fail over in seconds Seconds
WAN Management falls back to LTE; the carrier ticket is opened automatically Local inference and apps keep running
Disk Clustered NVMe keeps a replica on the other side None
In a reference incident from the console, power feed B dropped when a site breaker tripped. The appliance stayed up on feed A, AI Engine B workloads ran on A and the AI Factory for 22 minutes, with zero failed requests and p95 first-token latency peaking at 410 ms. Autonomous operations closed the investigation in 41 seconds.

Fleet operations

Every appliance is managed from the UPC Vektor console through the same five-stage lifecycle:
  • Enroll. Zero-touch: on first boot the appliance joins the console and pulls its policies.
  • Configure. Policies and blueprints by region, site or tenant.
  • Update. Firmware, models and patches roll out in waves: a canary site, then a region, then the fleet, with automatic pause on error.
  • Observe. Health, capacity, power and carbon for every site.
  • Remediate. Autonomous operations fix routine issues; engineers have remote console and power control for the rest.
Policy is inherited down a four-level hierarchy, global to region to site to appliance, and can be overridden at any level. A data-boundary policy set for EMEA, for example, applies to every appliance in that region automatically.
The built-in SRE Orchestrator correlates events into investigations, plans and runs the fix within guardrails, and asks a person to approve where policy requires. AI Coworker answers questions in plain language (“Why did DEN-Plant-02 page at 14:07?”), runs runbooks and updates tickets. Operation can be self-run, partner-managed or UnitedLayer-managed.

Security model

Control How it works
Data stays on site Models, prompts, outputs and embeddings never leave the appliance in sovereign mode
Encrypted at rest Hardware-sealed keys, secure boot and signed updates
Zero-trust access Your identity provider, role-based access and a full audit log
Isolated tenants Compute, storage, network and AI separated per business unit
Tamper-aware Lid, bezel and shock sensors raise alerts in the console
Recoverable Backups to an immutable cyber vault; DR in UnitedLayer facilities
Model supply chain Catalog models are security-scanned and license-reviewed; tenant policy controls which origins are allowed

Deployment and site readiness

The appliance ships pre-imaged, burned in and tested in a shock-rated crate, and goes live in a single visit.
  • 3U of rack space in a 19-inch rack at least 23 in deep
  • Two independent power circuits, about 1.4 kW peak in total
  • Two network handoffs per side (LAN and WAN), 25 GbE
  • Outbound connectivity to the UPC Vektor console, plus LTE coverage for backup management
  • Identity provider details for role-based access
  • Two or three first AI workloads and the VMs to consolidate, agreed in the sizing workshop
  • Data-boundary and model-origin policies for each tenant
To learn more about our services, contact us by clicking here!
↑