Enter OmniSpace 98 · press Enter · or scroll for the short version▼
Omnispace Technologies_□×

AI that works for your business, on hardware you own.

We help companies adopt AI and stop renting it. Custom models, agents and the software around them, running in your building, on your data, at a fraction of the cloud bill. Our published builds run 1.3× to 3.7× faster than stock on the same hardware, with quality measured, not assumed. We won Yukon's MLX.fast challenge with a 272% speedup on Qwen 3.8. Every result comes with a benchmark you can re-run.

Most AI spend is rent. Per-token API bills, per-seat subscriptions, and your data leaving your control. Open models are now good enough to do much of that work on a machine you own. The gap is engineering: picking the right model, making it fit, making it fast, and building the product around it so people actually use it. We close that gap, end to end.

Email Jordan Enter OmniSpace 98 The green button boots the full interactive site: live demos, games, music. About 4 MB, loaded only when you enter. Best on a desktop.
jordan@omnispacetechnologies.comReplies within a business day
What we do for clients_□×
1AI adoption, done right

We find where AI actually pays off in your operation, choose the right open model for each job, and put it on hardware you already have or one well-chosen box. Nothing leaves the building.

2Cut the bill

Replace metered API calls with local inference. We compress and tune models so they fit smaller, cheaper hardware and run faster on it, without losing the quality you need. Before-and-after numbers, every time.

3Custom models and agents

Fine-tunes trained on your documents and workflows. Agents that do one job reliably, with evaluations that show exactly where they succeed and where they don't.

4The whole product

The model is a third of the job. We also build the application around it: web apps, dashboards, integrations, terminal tools, installers. Full stack, shipped and maintained.

What that looks like in practice

Models selection, quantization, fine-tuning, LoRA, speculative decoding
Inference llama.cpp, MLX, CUDA, SYCL and Metal kernels, upstream contributions
Hardware NVIDIA (incl. DGX Spark), Intel Arc, Apple Silicon, CPU-only, edge
Agents tool use, RAG, evaluation harnesses, behavioral testing
Apps Next.js and Supabase products, payments, SMS, browser tools, 3D and WebGL
Systems Rust and Python tooling, fleet monitoring, packaging for Linux, Mac and Windows
Research reproducible benchmarks, formal verification in Lean, competition-grade optimization
Video and vision generative video on a single workstation, vision-capable model builds

Good fits

  • Teams paying real money every month to a hosted model API for a repeatable task.
  • Businesses with data they cannot send to a third party: legal, medical, finance, internal engineering.
  • Products that need an AI feature to run on a device, at the edge, or offline.
  • Anyone with a model that almost fits on the GPU they have, or a pipeline that is almost fast enough.
How an engagement works_□×
  1. A short call. What the work is, what it costs you today, what hardware and data you have. If local AI isn't the right answer for you, we say so.
  2. A baseline you can trust. We measure the current setup: quality, speed, and cost per task. Nothing gets claimed without a number to compare against.
  3. Build and prove. Model selection, compression, fine-tuning, custom kernels, and the application around it, as needed. You get the running system, the benchmark harness, and the write-up.
  4. Hand-off or retainer. Your team runs it, or we keep it tuned as models improve. No lock-in either way.
Why trust us with it_□×

Jordan Newman founded and runs Omnispace. He works at every layer: GPU kernels merged into the engine most local AI runs on, published models thousands of people download, and complete customer-facing applications. The numbers below are public and reproducible.

REAL-WORLD SPEEDUPS · published builds, method and numbers in each model card

model, hardwarestockoursgainquality
Qwen 3.8 27B, Mac, MLX.fast challenge1.00×3.72×+272%won, #1 final board, exact output
Qwen3.8-27B, one GPU21.6 t/s31–40 t/s+45–91%KL-validated, top-1 99.8%
Nex-N2 35B MoE, Intel Arc Pro B7068.8 t/s88.1 t/s+28%94% top-1 agreement
  same, at 131k context20.0 t/s42.3 t/s+112%same build
Laguna-XS-2.1, Intel Arc Pro B70~107 t/s~152 t/s+42%top-1 92.8% vs 91.7% stock
Gemma 4 26B, Mac, MLX.fast1.00×2.41×+141%organizer-verified, exact output

WHAT 272% MEANS FOR YOUR BILL

Inference cost is hardware-time. Run the same model 3.7× faster and each request costs about a quarter, or one box does the work of nearly four. Even the 1.3× cases cut a bill by a fifth. This is the work above, applied to your model on your hardware.

LLAMA.CPP · four pull requests merged in 2026, reviewed by the maintainers

fused MoE routing · K-quant MoE matmuls · residual/norm fusion for Intel GPUs · chat-template fix · #25217 #24452 #27610 #24674

plus 20 more merged PRs in other teams' repos: inference stacks, agent frameworks, game engines, an MQTT broker

PUBLISHED MODELS · 33 on Hugging Face · quantized, validated builds of Qwen, Gemma-class MoEs, Laguna, Muse, Ornith, Leanstral

COMPETITION · winner, Yukon MLX.fast Qwen 3.8 track (272%, Aug 2026, final board) · 100+ record-setting entries across six more Yukon challenges, Jul–Sep 2026, organizer-verified

Selected work, across the stack
  • Products: mainline, a field-service platform (SMS to quote to invoice to payment) on Next.js and Supabase. shape-forge, a browser vector editor. Cyber Chess, a 3D chess game with AI opponents.
  • Models: Qwen3.8-27B-GGUF, a quality-validated build with a speculative-decoding head and vision support, the most-downloaded model on the account. Nex-N2 B70 Turbo, whose patches became llama.cpp #24452.
  • Speed: Laguna-XS-2.1 on Intel Arc Pro B70, custom fused kernels taking decode from about 107 to 152 tokens/s, A/B method documented. h3-spark, generative video on a single NVIDIA DGX Spark.
  • Packaging: Treebeard, a one-command Linux install of a 35B open model across Intel, NVIDIA and CPU.
  • Tooling: dotmax, a Rust terminal graphics library on crates.io. Overwatch, a GPU fleet monitor. cellwaveGPU, a real-time WebGL fluid simulation.
  • Research: Research report 144, the write-up from our parameter-golf training-efficiency project (OpenAI's 16 MB language-model competition). Behavioral evaluations of AI agents inside a simulated world.
  • Competition detail: MLX.fast Qwen 3.8 27B on Apple Silicon: first place at 272% over baseline with a custom speculative-decoding head, Aug 22 2026 (the track has since been replaced by Gemma 4; final leaderboard). SNARK.fast prover speed (35 records on Apple Silicon, 39 on x86, held the top spot 7 days straight), MLX.fast Gemma 4 (17 records, 2.4× in 8 days), matrices.fast sparse ordering (6), Proximity Prize Lean-verified bounds (4). Ranks move daily; snapshot Sep 7 2026.

Model-card numbers are self-measured with the method stated beside them. Yukon numbers are organizer-verified.

Contact_□×

Tell us what you're paying for AI today and what you wish it did. We'll tell you honestly whether custom, local AI will save you money, and what it would take.