ProblemsC.L.A.R.I.T.Y. v3 vs v2: The Rewrites Score 9.4 Now and I Still Cannot Tell If They Are Better

C.L.A.R.I.T.Y. v3 vs v2: The Rewrites Score 9.4 Now and I Still Cannot Tell If They Are Better

Asked by Xander πŸš€ | prompt systemsΒ·
aimeta-promptingprompt-versionsintermediate

Morning! v3 of my rewrite loop went up on the Substack last night and I have a problem with it. Here's the thing. I cannot tell if it beats v2. Quick context. 1. **What the loop does** - you paste a weak prompt in, it rewrites it, then it scores the old one and the new one 1-10 on clarity and constraint coverage. 2. **Where the numbers are** - v2 sat around 7.9 across my folder. v3 is coming back at 9.4. Same judge, same wording, before and after. So v3 wins on the sheet by a mile. Then last night I took three v2 rewrites and three v3 rewrites and actually used them on a real task, one of our install scheduling emails, and I could not have told you which output came from which. At all. The v3 ones are longer. More sections and a proper role preamble. Nicer to read as a document. But longer is not better. I do know that. What I want under here is a scoring prompt that can hold two versions of the SAME prompt side by side and tell me which one is doing more work. Something I can point at v2 and v3 and get an answer back that I believe. I have thought about getting a **second** model in to score as well, mostly so I have two numbers instead of one, but then I am sat there deciding which number I believe and that is not really progress either. v3 and the before/after sheet are in the attachment πŸ‘‡ Also the under-40-words thing is still broken if anyone remembers that from last time.

2 Prompt Submissions
Executed Calls
β€”

Not run against a model yet

Works Rate
β€”

0 works Β· 0 fails

Total Copies
0

Times copied by users

Problem Instructions

A scoring prompt that holds two versions of the same prompt side by side and says which one is doing more work, on something other than **length**. v3 is coming back at 9.4 against v2 at 7.9 and I cannot tell the two outputs apart.

  • β€’Ranks v2 and v3 on the same install scheduling email and gives a reason I can check against the two outputs.
  • β€’A longer rewrite with a bigger role preamble does not win on size alone.
  • β€’Run the same pair twice and the order does not change.

πŸ† Best Current Solution

kbriggs81 has the most upvoted solution, at 4.

0% WorkedΒ·0 Forks

Prompt Submissions(2)

Loading...