10xEngineers

AI Infrastructure Software Services

AI Workloads to Your Custom Hardware.
From Applications to Kernels.

From model optimization to compiler development to custom runtime kernels, we build the AI software stack that gets your workloads running efficiently on any hardware, standard or custom.

Vision, language, multi-modal model deployment, SDK development, application pipelines
Model compression, inference speed up, and more
Expertise across Frontend frameworks, MLIR and LLVM
llama.cpp, IREE, Triton, C/Asm kernels
YOUR SILICONSubstrate
Custom & standard hardware
Trusted by Engineering Teams

Built alongside teams shaping RISC-V cores, AI silicon, and the tools that verify them.

OpenHW Group OpenHW Group
Woodpecker Technologies Woodpecker Technologies
Ainekko Ainekko
Partner

Everything Between the Model and the Hardware

Deploying AI efficiently requires more than powerful hardware. It requires a stack built in the right order, from the application layer down to the kernels that touch your silicon.

ML Applications

Building the ML software stack that brings your vision, language, and multimodal models to custom hardware.

  • Model zoo maintenance, porting existing vision, LLM, and VLM models to custom hardware
  • Pipelines for multi-modality (vision, audio, language) inference
  • AI software tooling and customized workflows

Model Optimization

Hardware-aware quantization and compression that shrinks models without breaking accuracy budgets.

  • QuantX: a hardware-aware quantization and compression framework
  • Vision, LLM, and VLM quantization
  • Custom bit-width and datatype exploration
  • Ongoing roadmap for more efficient speculative decoding

Compiler Development

MLIR and LLVM based compilation that turns your trained models into code your silicon can run.

  • Baltoro: our MLIR-based AI compiler
  • MLIR and LLVM based compilation for RISC-V CPUs and NPUs
  • Support for vision, LLM, and VLM models
  • Progressive lowering from framework graph to hardware-specific code

Runtime & Kernels

Custom runtime and kernel development that gets every layer of the stack running at full speed on your hardware.

  • KernelX: our agentic kernel development tool for automated kernel generation
  • Contributions to open frameworks like IREE for RISC-V accelerators
  • Runtime integration for custom and standard accelerators

Who We Help

AI infrastructure software services serve two kinds of teams, whichever end of the model-to-silicon path you're starting from.

Deploying AI on Your Hardware

Teams building or shipping silicon who need AI models running efficiently on it, from compiler support to custom kernels.

Semiconductor Companies AI Hardware Startups Enterprise Deployment Teams

Teams Training & Optimizing Their Own Models

Research and ML teams who train and iterate on their own models, and need that work to translate cleanly onto real hardware.

Research Organizations ML Engineering Teams Model Optimization Teams

Expertise, Productized

QuantX and Baltoro package our stack expertise into standalone products, with KernelX joining the lineup.

Model Compression

QuantX

Hardware-aware quantization for AI inference IP architects. Validate numeric formats like MXFP, BFP, NVFP, and FP16+INT before tape-out.

View Product
Compilation

Baltoro

The RISC-V-first, MLIR-based AI compiler with a modular Frontend, MLIR Compiler and Runtime for standard and custom silicon.

View Product
Kernel Development

KernelX

Our agentic kernel development tool, automating kernel generation for custom and standard hardware targets.

Coming Soon

Not sure which layer of the stack you need?

Tell us about your hardware and your models. We'll help you scope the right engagement.

Talk to an Engineer

What Partners Say

"10xEngineers have made outstanding contributions to the OpenHW community since joining in 2021, playing a key role in advancing the functionality, verification, and compliance of CVA6 and CV-Wally."

Florian 'Flo' Wohlrab
Head of OpenHW Foundation

"10xEngineers developed an LLVM-based tool that enhances our post-silicon validation process, significantly reducing Error Detection Latency and improving test coverage."

Woodpecker Technologies
Engineering Partner

Frequently Asked Questions

We specialize in the software stack that connects trained AI models to custom and standard silicon, spanning model compression (QuantX), MLIR/LLVM compiler engineering (Baltoro), custom ML kernels, and RISC-V toolchain development.

Yes. Our compiler team builds custom MLIR passes, dialects, and progressive lowering pipelines for unique hardware targets. This is the engineering foundation behind Baltoro.

Yes. We perform hardware-aware quantization, pruning, and compression on LLMs and VLMs through QuantX, then compile and deploy the optimized model on your target hardware.

Yes. Alongside standard RISC-V targets, we regularly build compiler and runtime support for proprietary accelerators and domain-specific instruction sets where off-the-shelf tooling doesn't exist yet.

Yes. Both are productized and can be adopted independently, or combined with our engineering services for a fully custom compiler and quantization pipeline.

Semiconductor companies, AI hardware startups, enterprise AI teams, and research organizations optimizing or deploying their models on custom silicon.

Yes. We offer ongoing engineering support, performance tuning, and toolchain maintenance beyond initial delivery, scoped to your project's needs.

Build Better AI Infrastructure

Let's discuss your compiler, runtime, or AI deployment challenges.

Contact Our Engineers