Efficient model family
Phase-gated 1B, 7B, 30B, 120B, 400B, 800B, and 1T+ model path with 1.58-bit architecture as the core efficiency wedge.
Public company site | Phase 0A funding now
Nantiro is building a 1.58-bit ternary AI model family that targets large-model capability with sharply lower memory, energy, and serving cost.
Company
The company mission is simple: make powerful AI cheaper to train, cheaper to serve, easier to deploy privately, and easier to measure for energy impact.
Phase-gated 1B, 7B, 30B, 120B, 400B, 800B, and 1T+ model path with 1.58-bit architecture as the core efficiency wedge.
Prepared-data handoff, tokenizer checks, training launch, checkpointing, loss monitoring, recovery, evaluation, and export workflow.
Target users include developers, startups, enterprises, institutions, and teams that need useful AI without frontier-lab infrastructure budgets.
Technology
Dense frontier AI is powerful, but memory bandwidth, GPU supply, electricity, and cooling make it expensive. Nantiro attacks the cost layer directly.
Dense models commonly store weights in FP16/BF16. Nantiro targets ternary weights {-1, 0, +1}, which mathematically need log2(3), or about 1.58 bits per weight. The goal is lower weight memory and fewer expensive multiply-heavy operations at serving time.
For the same parameter count, packed ternary weights can be about 10.1× smaller than FP16 weights. Training still needs high-precision master weights and optimizer state; the largest benefit appears at inference.
Target weight compression
Carbon-linked financing only works after measurement. Nantiro will track tokens, hardware power, energy per useful output, baseline comparison, and third-party MRV readiness from Phase 0 onward.
Measure, report, verify
Training means the model reads tokenized data, predicts the next token, measures error, updates weights, saves checkpoints, and repeats billions of times. That requires H100 GPU hours, fast storage, data preparation, recovery automation, benchmark evaluation, and operators who can keep the run healthy.
Comparison
They are shown to explain the market pain. Nantiro’s job is to validate a lower-cost path phase by phase.
Roadmap
Each stage must produce a clean scorecard before the next stage receives larger capital.
| Phase | Model | Target tokens | GPUs | Purpose |
|---|---|---|---|---|
| 0A | Nantiro 1B | 20B-30B | 8× H100 | First 1.58-bit stack proof: data, tokenizer, training, checkpoint, dashboard, eval. |
| 0B | Nantiro 7B | 100B-150B | 64× H100 | Distributed training proof, LR/threshold validation, loss-recovery proof. |
| 0C | Nantiro 30B | 300B-500B | 128× H100 | System validation and planned open-source release after scorecards pass. |
| 1 | Nantiro 120B | 1.2T min / 2.4T target | 1,024× H100 | Foundation model attempt after Phase 0C proves the recipe. |
| 2 | Nantiro 400B MoE | 6T-8T | 2,048× H100 | MoE scale-up with ternary experts and compact serving goal. |
| 3 | Nantiro 800B MoE | 12T-16T | 4,096× H100 | Large MoE validation with frontier-scale inference ambition. |
| 4 | Nantiro 1T+ MoE | 20T+ | 4,096× H100 | Long-horizon trillion-class stage only after Phase 3 proof. |
Business model
These are planning scenarios using the current website pricing bands: enterprise ₹5L-₹20L/year, API ₹0.01 per 1K tokens, Chat Plus ₹199-₹499/month, and Edge Offline ₹999 one-time.
| Phase 0C scenario | Enterprise | API | Consumer | Possible annual income |
|---|---|---|---|---|
| Conservative | 20 pilots × ₹5L = ₹1.0Cr | 50B tokens/mo = ₹0.6Cr/yr | 25k Chat + 10k Edge = ₹9.97Cr | ~₹11.6Cr |
| Base | 50 customers × ₹12.5L = ₹6.25Cr | 250B tokens/mo = ₹3.0Cr/yr | 100k Chat + 50k Edge = ₹40.88Cr | ~₹50.1Cr |
| Upside | 150 customers × ₹20L = ₹30Cr | 1T tokens/mo = ₹12Cr/yr | 250k Chat + 150k Edge = ₹104.69Cr | ~₹146.7Cr |
These are commercial planning scenarios, not guaranteed revenue. Phase 0C must first pass quality, safety, reliability, and cost-per-token validation.
Ultra-low-cost tokens for developers, small teams, and startups.
On-prem or private-cloud model serving for teams that cannot send data to closed APIs.
After Phase 0C validation, the 30B release can build community trust and paid enterprise demand.
Funding
8× H100 for 28 days, 20B-30B tokens, storage, data checks, benchmark run, monitoring, checkpoint recovery, legal/IP setup, and contingency.
Carbon-linked financing
Nantiro can still build the measurement layer now, so energy efficiency becomes a financing and reporting advantage later.
Assumption: dense serving baseline 1.25Wh/query, Nantiro target 0.35Wh/query, saving 0.90Wh/query. At 1B queries/day, annual saving is 328.5GWh.
Using 0.70kg CO₂/kWh grid factor. This replaces the earlier 324k-ton card with a clearer formula and lower, more defensible estimate.
At ₹500-₹2,500 per verified ton, if a suitable methodology, registry pathway, and third-party verification are accepted.
Public engagement
Answer a short opinion survey. Based on your response, the site asks one follow-up and then shows how Nantiro addresses your concern.
Contact