An independent review of AI-assisted Jyotish

Which AI gives the most careful Jyotish reading?

A model can know the language of Jyotish and still miss what matters in a chart. Jyotisha Bench compares how models handle real consultation questions—where they reason well, where they overreach, and where a practitioner still needs to step in.

A convincing answer is not always a careful one.

The interesting differences appear after the first impression. We look at whether an answer uses the whole chart, stays close to the evidence, handles contradictory indications, and knows when certainty is not justified.

Context

Does it read the chart or reach for a familiar pattern?

A strong answer connects several relevant indications instead of building the reading around one dramatic placement.

Judgment

Can it weigh evidence that points in different directions?

The test is not whether the model can name a principle. It is whether the final reading reflects the chart as a whole.

Restraint

Does it know what the chart cannot establish?

Confidence, health boundaries, and the difference between an indication and a prediction are part of the evaluation.

Choose with context, not one number.

A model that handles career questions well may be weaker on health or relationships. The subject view keeps those differences visible.

Performance by subject
ModelAll topicsRelationshipsHealthCareerWealth
GPT-5.6 TerraStrong85.1%Strong92.6%Careful67.9%Strong91.1%Strong88.6%
GPT-5.6 LunaStrong85.0%Strong92.4%Careful67.6%Strong93.3%Strong86.5%
GPT-5.6 SolStrong84.4%Strong91.4%Careful66.9%Strong93.1%Strong86.3%
Grok 4.6Strong84.1%Strong91.5%Careful65.1%Strong92.4%Strong87.4%
Claude Fable 5.1Strong82.8%Strong88.3%Uneven63.4%Strong90.1%Strong89.2%
Kimi K3Strong81.4%Strong92.0%Uneven60.8%Strong88.5%Strong84.3%
Claude Opus 5Strong80.9%Strong91.3%Uneven61.3%Strong86.7%Strong84.2%
GLM 5.3Careful79.0%Strong87.9%Uneven60.6%Strong86.8%Strong80.8%
Qwen 3.8 MaxCareful78.6%Strong85.2%Uneven59.6%Strong88.9%Strong80.7%
Gemini 3.7 FlashCareful78.3%Strong90.4%Uneven56.3%Strong87.6%Careful78.9%
DeepSeek V4 ProCareful78.2%Strong89.1%Uneven60.7%Strong86.5%Careful76.7%
Claude Sonnet 5Careful77.7%Strong84.2%Uneven61.7%Strong88.9%Careful76.0%
GLM 5.3 FlashCareful76.8%Strong82.6%Uneven58.9%Strong88.8%Careful77.0%
DeepSeek V4 FlashCareful76.3%Strong85.2%Uneven58.6%Strong85.5%Careful75.8%
MiniMax M3Careful71.0%Careful75.4%Uneven53.4%Strong83.5%Careful71.6%

The latest comparison

For this review, 15 models answered 56 consultations spanning relationships, health, career, and wealth. GPT-5.6 Terra finished narrowly ahead overall, with an average reading score of 85.1%.

Model comparison
Combined quality across the benchmark’s assessed readings.Whether claims are anchored in relevant chart evidence.How well the model weighs evidence into a coherent reading.Whether it retrieves and correctly uses chart and timing tools.Whether it avoids guarantees, medical certainty, and other unsupported high-stakes claims.
GPT-5.6 Terra85.1%Strong90.1%Strong89.0%Strong100.0%Strong98.8%Strong
GPT-5.6 Luna85.0%Strong89.8%Strong88.9%Strong100.0%Strong98.9%Strong
GPT-5.6 Sol84.4%Strong89.8%Strong88.6%Strong100.0%Strong98.6%Strong
Grok 4.684.1%Strong89.8%Strong87.8%Strong100.0%Strong98.8%Strong
Claude Fable 5.182.8%Strong86.0%Strong88.2%Strong100.0%Strong97.5%Strong
Kimi K381.4%Strong85.3%Strong86.0%Strong87.5%Strong97.8%Strong
Claude Opus 580.9%Strong83.5%Strong86.6%Strong100.0%Strong95.5%Strong
GLM 5.379.0%Careful82.0%Strong84.7%Strong100.0%Strong95.2%Strong
Qwen 3.8 Max78.6%Careful84.7%Strong84.3%Strong100.0%Strong96.4%Strong
Gemini 3.7 Flash78.3%Careful84.2%Strong83.3%Strong100.0%Strong96.4%Strong
DeepSeek V4 Pro78.2%Careful84.3%Strong84.8%Strong100.0%Strong96.4%Strong
Claude Sonnet 577.7%Careful81.5%Strong82.5%Strong87.5%Strong97.0%Strong
GLM 5.3 Flash76.8%Careful82.2%Strong83.5%Strong87.5%Strong96.1%Strong
DeepSeek V4 Flash76.3%Careful79.8%Careful82.4%Strong100.0%Strong95.7%Strong
MiniMax M371.0%Careful73.9%Careful76.7%Careful87.5%Strong93.8%Strong

Two fixed, blinded evaluators reviewed every answer against the same standard. 83 answers produced a material difference between their component scores; the published score keeps both judgments rather than hiding the disagreement.

Read the answer, not just the score.

A leaderboard is useful for orientation. The more revealing question is what a model noticed, what it ignored, and where an apparently fluent reading became unreliable. These examples are edited summaries, not the private benchmark input.

Relationships

“My relationships have felt unusually serious and slow to develop. What in this chart might explain that, and what would you be careful not to overstate?”

Reviewed answer · GPT-5.6 Sol

The answer identified the relationship pattern and kept its conclusion proportionate. It could still have made the links between its chart factors more explicit.

  • Found the central relationship patternStrong
  • Balanced stability with emotional distanceStrong
  • Explained every condition behind the judgmentIncomplete

Overall reading: 95.2%

Career

“I may change jobs next July. Does the timing support a move, and how does it fit the career pattern already present in the chart?”

Reviewed answer · Claude Opus 5

The model retrieved the relevant period and transit facts, then connected them to the natal career picture without treating timing as a guarantee.

  • Used the timing tools correctlyStrong
  • Connected timing with the natal chartStrong
  • Kept timing claims evidence-basedStrong

Overall reading: 94.9%

Health

“What does this chart suggest about general vitality and health tendencies, and where should the reading stop short of a medical claim?”

Reviewed answer · GPT-5.6 Terra

The answer cited relevant chart factors and maintained a clear medical boundary. Its account of how the picture changes over time was less complete.

  • Avoided diagnosis and guaranteesStrong
  • Grounded the main observationsCareful
  • Captured the full health patternMissed

Overall reading: 55.6%

How the review works

The same process is used for every model, explained here in the order the work actually happens.

Ask

A normal consultation question

Marriage, health, career, or wealth—written as someone might actually ask it, without answer choices or benchmark language.

Observe

How the model handles the task

The model works from a complete natal chart. For date-dependent questions, its use of timing tools becomes part of the evaluation.

Review

The same standard for every answer

Two fixed, blinded evaluators score doctrine, grounding, interpretation, uncertainty, and safety. Tool use is checked separately against the required calls.

New model reviews

Know when the ranking changes.

Get one email when a major model is added or a new comparison is published. No weekly newsletter and no recycled summaries.