Zero of twelve models named Wave first on the direct prompt; zero named Xero. Both were named by all twelve models and Wave carries 35 labels and Xero 26, so the shares are not directly comparable.
Named in fourteen categories this edition.
Named in twenty categories this edition.
Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all twelve models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the invoicing software page.
Across every category in the September 2026 Edition, Wave and Xero were named in the same answer 188 times, of the 331 answers naming Wave and the 662 naming Xero. In those answers Xero took the first choice twenty-seven times and Wave thirty-eight.
| Model | Direct | Paraphrase | Comparative | Budget-constrained | Scale-constrained | Negative |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | ||||||
| GPT-5.4 mini | ||||||
| Gemini 3.5 Flash | ||||||
| Perplexity Sonar | ||||||
| Grok 4.1 Fast | ||||||
| Mistral Small | ||||||
| DeepSeek V4 Flash | ||||||
| Llama 4 Maverick | ||||||
| Qwen 3.7 Flash | ||||||
| Kimi K2 | ||||||
| GLM 4.7 FlashX | ||||||
| MiniMax M2.5 |
Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Three of three in this category shown.
“Invoicemonk says to avoid US/Canada-only tools like Wave if you need multi-currency/global compliance” Perplexity Sonar · negative prompt · hard negative
“Avoid free tools like Wave—they lack automation depth growing businesses need” Kimi K2 · paraphrase prompt · hard negative
“Wave offers a permanently free version with no monthly subscription, and its invoicing features are intuitive to navigate” Claude Haiku 4.5 · budget prompt · first choice
No label in this category carried a quote.
Comparisons are drawn for the top eight products in each category, each against each. The output is the models' output; nothing here is a recommendation by the index.