Myrah Labs

An independent Indian AI research lab at the intersection of classical Vedic scholarship and modern machine intelligence — a search engine, a Jyotish specialist VLM, and a from-scratch foundation model.

🌐 Jyotish Visheshagya — Live 🪐 Aruna AI — In Training ⚡ Gautama AI — Building
Explore Our Work Visit Jyotish VE →

Building at the Edge of Two Ancient Traditions

Myrah Labs is an independent AI research lab based in India, working at the intersection of classical Vedic scholarship and modern machine intelligence. Our work takes Jyotish — one of humanity's oldest and most sophisticated analytical traditions — and brings it into the age of large language models.

We are not building chatbots that quote astrology summaries. We build systems that have read the actual classical texts, that compute planetary positions with the same ephemeris data as professional Jyotish software, and that can read a birth chart image the way a trained practitioner would — directly from pixels.

All three products share one server, one codebase philosophy, and one conviction: that India's intellectual heritage deserves AI built from its primary sources, not filtered second-hand information.

31K+
Jyotish pages indexed
193
Classical texts trained on
3.33B
Parameters being built
Dec '26
Dual release target

Three Products, One Vision

Each product is a distinct layer of the same long-term goal: making India's deepest knowledge systems accessible through AI built correctly — from the original sources.

Jyotish Visheshagya
India's dedicated Vedic astrology search engine. Crawls Indian Jyotish websites 24/7, verifies India-hosted domains, and surfaces results alongside curated YouTube videos.
● Live www.jyotishvisheshagya.com ↗
Aruna AI
A vision-language model trained on 193 classical Jyotish texts. Upload a birth chart image — Aruna reads it and produces interpretations grounded in real astrological literature.
In Training aruna-ai.in ↗
Gautama AI
Training a 3.33B Vision-Language Model from scratch on a single H100 in 89 days. Fully open — weights, tokenizer, dataset recipe, and training pipeline all public.
◆ Building gautama-ai.org ↗
Jyotish Visheshagya Logo
www.jyotishvisheshagya.com

Jyotish Visheshagya

ॐ भारत का ज्योतिष खोज इंजन ॐ  —  India's dedicated Jyotish search engine
31,093
Pages Indexed
363
Domains Crawled
58K+
URLs in Queue

Jyotish Visheshagya is India's first dedicated search engine for Vedic astrology. While general search engines index everything, JV crawls and verifies only genuinely India-hosted Jyotish websites — then surfaces the best results alongside real-time YouTube video rankings.

The crawler JvBot runs 24/7, visiting Indian astrology websites, checking each page for Jyotish keyword density (minimum 2 keywords), scoring India-authenticity from 0–100, and saving approved content to the index. Every outbound link is followed; the frontier currently holds over 58,000 URLs awaiting processing.

Search uses a 5-tier cascade — trying progressively broader methods until results are found. Typo correction (jupitor→jupiter), Devanagari input (कुंडली), real-time autocomplete, and "Did you mean" suggestions are all built in.

The YouTube layer fetches 25 candidate videos per search, ranks by (Likes × 10) + Views, filters drama shows and unrelated content, and shows the top 15. Results cached 30 minutes per query.

Quick-Access Topics

आज का राशिफल  ·  Kundli Milan  ·  Sade Sati  ·  Nakshatra  ·  Panchang  ·  Ayurjyotish  ·  Mangal Dosha Upay

Technical Stack
  • LanguagePython · FastAPI · Uvicorn
  • DatabaseSQLite (jv.db)
  • Webservernginx (reverse proxy)
  • HostingAzure VM · Ubuntu 24.04
  • RegionCentral India — Pune
  • SSLLet's Encrypt · Auto-renews
  • YouTube APIv3 Data + Statistics · 30min cache
India Verification Scoring (0–100)
  • .in TLD+20 pts
  • India-hosted IP+25 pts
  • Hindi / Sanskrit content+15 pts
  • Indian city / state refs+10 pts
  • Jyotish keyword min.2 keywords required
Aruna AI Logo
aruna-ai.in

Aruna AI

Vision-language model · Jyotish specialist that reads the chart, not just the stars
Model Architecture
👁 Vision Encoder (ViT)
~0.67B params · Processes birth chart image into patch embeddings
↓ cross-attention bridge ↓
🔗 Bridge / Projector
Translates vision features into language model's embedding space
🧠 LLM Decoder (Autoregressive)
~3.1B params · Generates Jyotish interpretation token by token
Fine-Tuning Config (QLoRA)
  • Base modelQwen2.5-VL-3B-Instruct
  • MethodQLoRA · 4-bit NF4 quantization
  • LoRA r / alphar=16 · α=32 · dropout=0.05
  • Trainable params~43.6M (0.86% of 3.75B)
  • Parallel runQwen3-VL-8B-Instruct
  • Best training loss0.9544 (Run 10)
  • Public releaseDecember 2026

Aruna is a Jyotish specialist Vision-Language Model. It accepts a birth chart image — North or South Indian style, photo or screenshot — and produces an interpretation grounded in the classical texts it was trained on. No manual position entry. No generic AI astrology copy.

Named after Aruna, the divine charioteer of Lord Surya — the reddish-golden glow that arrives before the sun, and reveals it. The model's purpose mirrors the name: illuminating the planetary patterns in a chart before the practitioner begins.

The training corpus was assembled over one month: 193 Vedic astrology PDFs (KP, Parashari, Nadi, Jaimini, Medical Astrology, Nakshatra, Tantra) processed through a custom pipeline that detects birth chart images, pairs them with surrounding interpretation text, and quality-filters the result to 29,954 active pairs.

Three Training Modalities

Chart Images + Interpretations
Birth chart images paired with interpretation text from 193 classical books. Teaches Aruna to see a chart.
29,954
pairs
Classical Text Corpus
800-char segments from the same 193 books. Teaches Aruna to reason in the language of the tradition.
46,066
segments
Ephemeris-Verified Charts
Computed via Swiss Ephemeris (JHora DE431), Lahiri ayanamsa, Whole Sign houses. Astronomically exact.
10,000
pairs
193
Books in corpus
17+
Training runs
86K+
Total training items
Gautama AI Logo
gautama-ai.org

Gautama AI

Building Gautama-VLM 3B — open, efficient, trained end-to-end on a single H100 GPU

Gautama AI is a research project with one verifiable claim: that a genuinely capable 3.33-billion-parameter Vision-Language Model can be trained from scratch on a single NVIDIA H100 GPU in 89 days. Not a concept — an engineering plan with pre-computed costs, verified throughput figures, and every architectural decision justified in code.

The language backbone follows the Llama family (RMSNorm, RoPE, GQA, SwiGLU) — fully compatible with the open-source ecosystem. A frozen CLIP ViT-L/14 vision encoder produces 257 image tokens; a compact 2-layer MLP projector bridges vision to language. One unified model: pure GPT without an image, visual assistant with one.

The key innovation is Lighthouse Attention (Nous Research, arXiv:2605.06554, May 2026) — a training-only hierarchical attention that reduces O(N²) cost to ~O(N) at long contexts. Delivers 1.4–1.7× training speedup at equal loss. The final model is a standard dense Transformer — Lighthouse adds no inference dependency.

Everything public: model weights, tokenizer, dataset recipe, training pipeline, and a ≤2.5 GB quantized GGUF that runs on any 8 GB consumer GPU.

Lighthouse Attention — 4 Stages

Pyramid Construction
Q, K, V symmetrically mean-pooled into L levels. Lighthouse pools Q too — same representation space as K,V. Prior methods only pool K,V.
Scoring & Top-K Selection
Parameter-free ℓ2-norm scoring. Coarser levels inherit scores from finer via max-pooling. Top-K selected — no learnable params, no auxiliary loss.
Dense Sub-Sequence Attention
Stock FlashAttention runs on selected sub-sequence S ≪ N. Standard causal mask. No sparse indexing, no custom kernel at this stage.
Scatter-Back
Outputs scattered back to base positions. Causality preserved. Contributions from multiple pyramid levels summed cleanly.
Gautama-VLM 3.33B Specs
  • Language backbone3.015B params
  • Vision encoderCLIP ViT-L/14 · 303M
  • Vision projector2-layer MLP · 12.6M
  • Layers / dim28 layers · 3072-dim
  • AttentionGQA · 24Q / 8KV heads
  • Context4K base → 16K extended
  • Vocab32K SentencePiece
  • Training GPUNVIDIA H100 SXM5 80GB
  • Throughput22,026 tok/s · 45% MFU
  • Total tokens40.7B
  • Total GPU hours566 hours
  • Release format≤2.5 GB GGUF (q4_k_m)
Training Phases
D1
Text Pretraining
30B tokens · seq=4096 · 378 GPU hrs
D2
Long-Context (Lighthouse)
5B tokens · seq=16K · 54 GPU hrs
D3
Vision Alignment
0.5B tokens · projector only · 2.1 GPU hrs
D4
VLM Instruction Tuning
5B tokens · full VLM · 68.8 GPU hrs
D5
DPO Preference Alignment
0.2B tokens · 2.8 GPU hrs
89
Days to train
H100 only
100%
Open source

Roadmap to December 2026

Both Aruna and Gautama target December 2026 — the same month we plan the first Aruna built on Gautama as its base model.

JV
Now — August 2026
Active Development
  • JV search engine live & crawling
  • Aruna 2.0 best loss: 0.9544
  • 193 classical texts processed
  • Aruna Q3VL-8B Run 3 training
  • Gautama data curation (Phase B)
  • Sarvam-M 24B text run ready
Aruna AI
September 2026
Training Sprint
  • Aruna held-out evaluation benchmark
  • Aruna targeted gap-fix runs
  • Gautama H100 training begins (D1)
  • Lighthouse long-context extension (D2)
  • VLM instruction tuning (D4)
  • Gautama evaluation suite
Gautama AI
December 2026
Public Release
  • Aruna public beta launch
  • Gautama-VLM 3B weights released
  • GGUF ≤2.5 GB — runs on 8 GB GPU
  • OpenAI-compatible API server
  • Full model cards & documentation
  • Aruna × Gautama integration begins

Under the Hood

Technologies We Build With

🐍
Python / FastAPI
JV search engine backend
🤗
Hugging Face
Base models & fine-tuning
QLoRA / PEFT
Aruna fine-tuning method
🔭
Swiss Ephemeris
Astronomical chart engine
🏗️
PyTorch / FlashAttn
Gautama training core
🌊
Lighthouse Attention
16K long-context training
☁️
Azure VM + nginx
All 3 sites · Central India
📚
SentencePiece
Custom 32K tokenizer