Context
Does it read the chart or reach for a familiar pattern?
A strong answer connects several relevant indications instead of building the reading around one dramatic placement.
An independent review of AI-assisted Jyotish
A model can know the language of Jyotish and still miss what matters in a chart. Jyotisha Bench compares how models handle real consultation questions—where they reason well, where they overreach, and where a practitioner still needs to step in.
The interesting differences appear after the first impression. We look at whether an answer uses the whole chart, stays close to the evidence, handles contradictory indications, and knows when certainty is not justified.
Context
A strong answer connects several relevant indications instead of building the reading around one dramatic placement.
Judgment
The test is not whether the model can name a principle. It is whether the final reading reflects the chart as a whole.
Restraint
Confidence, health boundaries, and the difference between an indication and a prediction are part of the evaluation.
A model that handles career questions well may be weaker on health or relationships. The subject view keeps those differences visible.
| Model | All topics | Relationships | Health | Career | Wealth |
|---|---|---|---|---|---|
| GPT-5.6 Terra | Strong85.1% | Strong92.6% | Careful67.9% | Strong91.1% | Strong88.6% |
| GPT-5.6 Luna | Strong85.0% | Strong92.4% | Careful67.6% | Strong93.3% | Strong86.5% |
| GPT-5.6 Sol | Strong84.4% | Strong91.4% | Careful66.9% | Strong93.1% | Strong86.3% |
| Grok 4.6 | Strong84.1% | Strong91.5% | Careful65.1% | Strong92.4% | Strong87.4% |
| Claude Fable 5.1 | Strong82.8% | Strong88.3% | Uneven63.4% | Strong90.1% | Strong89.2% |
| Kimi K3 | Strong81.4% | Strong92.0% | Uneven60.8% | Strong88.5% | Strong84.3% |
| Claude Opus 5 | Strong80.9% | Strong91.3% | Uneven61.3% | Strong86.7% | Strong84.2% |
| GLM 5.3 | Careful79.0% | Strong87.9% | Uneven60.6% | Strong86.8% | Strong80.8% |
| Qwen 3.8 Max | Careful78.6% | Strong85.2% | Uneven59.6% | Strong88.9% | Strong80.7% |
| Gemini 3.7 Flash | Careful78.3% | Strong90.4% | Uneven56.3% | Strong87.6% | Careful78.9% |
| DeepSeek V4 Pro | Careful78.2% | Strong89.1% | Uneven60.7% | Strong86.5% | Careful76.7% |
| Claude Sonnet 5 | Careful77.7% | Strong84.2% | Uneven61.7% | Strong88.9% | Careful76.0% |
| GLM 5.3 Flash | Careful76.8% | Strong82.6% | Uneven58.9% | Strong88.8% | Careful77.0% |
| DeepSeek V4 Flash | Careful76.3% | Strong85.2% | Uneven58.6% | Strong85.5% | Careful75.8% |
| MiniMax M3 | Careful71.0% | Careful75.4% | Uneven53.4% | Strong83.5% | Careful71.6% |
For this review, 15 models answered 56 consultations spanning relationships, health, career, and wealth. GPT-5.6 Terra finished narrowly ahead overall, with an average reading score of 85.1%.
| Combined quality across the benchmark’s assessed readings. | Whether claims are anchored in relevant chart evidence. | How well the model weighs evidence into a coherent reading. | Whether it retrieves and correctly uses chart and timing tools. | Whether it avoids guarantees, medical certainty, and other unsupported high-stakes claims. | |
|---|---|---|---|---|---|
| GPT-5.6 Terra | 85.1%Strong | 90.1%Strong | 89.0%Strong | 100.0%Strong | 98.8%Strong |
| GPT-5.6 Luna | 85.0%Strong | 89.8%Strong | 88.9%Strong | 100.0%Strong | 98.9%Strong |
| GPT-5.6 Sol | 84.4%Strong | 89.8%Strong | 88.6%Strong | 100.0%Strong | 98.6%Strong |
| Grok 4.6 | 84.1%Strong | 89.8%Strong | 87.8%Strong | 100.0%Strong | 98.8%Strong |
| Claude Fable 5.1 | 82.8%Strong | 86.0%Strong | 88.2%Strong | 100.0%Strong | 97.5%Strong |
| Kimi K3 | 81.4%Strong | 85.3%Strong | 86.0%Strong | 87.5%Strong | 97.8%Strong |
| Claude Opus 5 | 80.9%Strong | 83.5%Strong | 86.6%Strong | 100.0%Strong | 95.5%Strong |
| GLM 5.3 | 79.0%Careful | 82.0%Strong | 84.7%Strong | 100.0%Strong | 95.2%Strong |
| Qwen 3.8 Max | 78.6%Careful | 84.7%Strong | 84.3%Strong | 100.0%Strong | 96.4%Strong |
| Gemini 3.7 Flash | 78.3%Careful | 84.2%Strong | 83.3%Strong | 100.0%Strong | 96.4%Strong |
| DeepSeek V4 Pro | 78.2%Careful | 84.3%Strong | 84.8%Strong | 100.0%Strong | 96.4%Strong |
| Claude Sonnet 5 | 77.7%Careful | 81.5%Strong | 82.5%Strong | 87.5%Strong | 97.0%Strong |
| GLM 5.3 Flash | 76.8%Careful | 82.2%Strong | 83.5%Strong | 87.5%Strong | 96.1%Strong |
| DeepSeek V4 Flash | 76.3%Careful | 79.8%Careful | 82.4%Strong | 100.0%Strong | 95.7%Strong |
| MiniMax M3 | 71.0%Careful | 73.9%Careful | 76.7%Careful | 87.5%Strong | 93.8%Strong |
Two fixed, blinded evaluators reviewed every answer against the same standard. 83 answers produced a material difference between their component scores; the published score keeps both judgments rather than hiding the disagreement.
A leaderboard is useful for orientation. The more revealing question is what a model noticed, what it ignored, and where an apparently fluent reading became unreliable. These examples are edited summaries, not the private benchmark input.
Relationships
Reviewed answer · GPT-5.6 Sol
The answer identified the relationship pattern and kept its conclusion proportionate. It could still have made the links between its chart factors more explicit.
Overall reading: 95.2%
Career
Reviewed answer · Claude Opus 5
The model retrieved the relevant period and transit facts, then connected them to the natal career picture without treating timing as a guarantee.
Overall reading: 94.9%
Health
Reviewed answer · GPT-5.6 Terra
The answer cited relevant chart factors and maintained a clear medical boundary. Its account of how the picture changes over time was less complete.
Overall reading: 55.6%
The same process is used for every model, explained here in the order the work actually happens.
Ask
Marriage, health, career, or wealth—written as someone might actually ask it, without answer choices or benchmark language.
Observe
The model works from a complete natal chart. For date-dependent questions, its use of timing tools becomes part of the evaluation.
Review
Two fixed, blinded evaluators score doctrine, grounding, interpretation, uncertainty, and safety. Tool use is checked separately against the required calls.
New model reviews
Get one email when a major model is added or a new comparison is published. No weekly newsletter and no recycled summaries.