Eight of twelve models named Prophix first on the direct prompt; one named Sage Intacct. Prophix was named by eleven of the twelve models and Sage Intacct by ten and Prophix carries 30 labels and Sage Intacct 14, so the shares are not directly comparable.
Named in nine categories this edition.
Named in seventeen categories this edition.
Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all twelve models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the financial consolidation page.
Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Four of four in this category shown.
“Tools like OneStream or Prophix are powerful but geared toward larger/enterprise (higher cost, 6+ mo impl.)” Grok 4.1 Fast · paraphrase prompt · soft negative
“Prophix One is particularly notable for its ability to connect consolidation to the close and planning in one platform.” Llama 4 Maverick · direct prompt · first choice
“Choose Prophix if you want the best balance of consolidation depth, planning integration, and mid-market fit” Kimi K2 · direct prompt · first choice
“Prophix One – best if you want close, planning, and consolidation in one platform with strong governance.” GLM 4.7 FlashX · direct prompt · first choice
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Two of two in this category shown.
“Avoid spreadsheets or full ERPs like NetSuite/Sage Intacct—they're error-prone or too expensive” Grok 4.1 Fast · budget prompt · hard negative
“avoid ... heavy ERP replacements (like NetSuite or Sage Intacct)” Gemini 3.5 Flash · budget prompt · hard negative
Comparisons are drawn for the top three products in each category. The output is the models' output; nothing here is a recommendation by the index.