AI Workloads to Your Custom Hardware.
From Applications to Kernels.
From model optimization to compiler development to custom runtime kernels, we build the AI software stack that gets your workloads running efficiently on any hardware, standard or custom.
Everything Between the Model and the Hardware
Deploying AI efficiently requires more than powerful hardware. It requires a stack built in the right order, from the application layer down to the kernels that touch your silicon.
ML Applications
Building the ML software stack that brings your vision, language, and multimodal models to custom hardware.
- Model zoo maintenance, porting existing vision, LLM, and VLM models to custom hardware
- Pipelines for multi-modality (vision, audio, language) inference
- AI software tooling and customized workflows
Model Optimization
Hardware-aware quantization and compression that shrinks models without breaking accuracy budgets.
- QuantX: a hardware-aware quantization and compression framework
- Vision, LLM, and VLM quantization
- Custom bit-width and datatype exploration
- Ongoing roadmap for more efficient speculative decoding
Compiler Development
MLIR and LLVM based compilation that turns your trained models into code your silicon can run.
- Baltoro: our MLIR-based AI compiler
- MLIR and LLVM based compilation for RISC-V CPUs and NPUs
- Support for vision, LLM, and VLM models
- Progressive lowering from framework graph to hardware-specific code
Runtime & Kernels
Custom runtime and kernel development that gets every layer of the stack running at full speed on your hardware.
- KernelX: our agentic kernel development tool for automated kernel generation
- Contributions to open frameworks like IREE for RISC-V accelerators
- Runtime integration for custom and standard accelerators
Who We Help
AI infrastructure software services serve two kinds of teams, whichever end of the model-to-silicon path you're starting from.
Deploying AI on Your Hardware
Teams building or shipping silicon who need AI models running efficiently on it, from compiler support to custom kernels.
Teams Training & Optimizing Their Own Models
Research and ML teams who train and iterate on their own models, and need that work to translate cleanly onto real hardware.
Expertise, Productized
QuantX and Baltoro package our stack expertise into standalone products, with KernelX joining the lineup.
QuantX
Hardware-aware quantization for AI inference IP architects. Validate numeric formats like MXFP, BFP, NVFP, and FP16+INT before tape-out.
View ProductBaltoro
The RISC-V-first, MLIR-based AI compiler with a modular Frontend, MLIR Compiler and Runtime for standard and custom silicon.
View ProductKernelX
Our agentic kernel development tool, automating kernel generation for custom and standard hardware targets.
Coming SoonNot sure which layer of the stack you need?
Tell us about your hardware and your models. We'll help you scope the right engagement.
What Partners Say
"10xEngineers have made outstanding contributions to the OpenHW community since joining in 2021, playing a key role in advancing the functionality, verification, and compliance of CVA6 and CV-Wally."
"10xEngineers developed an LLVM-based tool that enhances our post-silicon validation process, significantly reducing Error Detection Latency and improving test coverage."
Frequently Asked Questions
Yes. Our compiler team builds custom MLIR passes, dialects, and progressive lowering pipelines for unique hardware targets. This is the engineering foundation behind Baltoro.
Yes. We perform hardware-aware quantization, pruning, and compression on LLMs and VLMs through QuantX, then compile and deploy the optimized model on your target hardware.
Yes. Alongside standard RISC-V targets, we regularly build compiler and runtime support for proprietary accelerators and domain-specific instruction sets where off-the-shelf tooling doesn't exist yet.
Yes. Both are productized and can be adopted independently, or combined with our engineering services for a fully custom compiler and quantization pipeline.
Semiconductor companies, AI hardware startups, enterprise AI teams, and research organizations optimizing or deploying their models on custom silicon.
Yes. We offer ongoing engineering support, performance tuning, and toolchain maintenance beyond initial delivery, scoped to your project's needs.
Build Better AI Infrastructure
Let's discuss your compiler, runtime, or AI deployment challenges.
Contact Our Engineers