Flag mishaps in the analysis to feed benchmarks
Enable users to flag any part of an analysis that was performed but shows a misinterpretation from the LLM and requires an adjustment.
Allows to fine tune the analysis by feeding the example to the benchmark samples, which allows to build your own verified dataset over time, and will give you tools to adjust the analysis pipeline to improve its quality.
Status: Planned
Log in to comment and vote
No comments yet
Be the first to share your thoughts.