controlbun
Extracting a steering direction takes about a day. Deciding whether to trust it takes weeks, and none of that work carries over to anybody else.
I have a for and I .
Then what you want is not our opinion, it is somebody else's check. A submission here shows what its author measured, what they did not, and anything a third party found afterwards. Nothing is scored and nothing is ranked. The three pro-human directions on this site come from one dataset at one layer and differ only in the estimator, every check their author ran says all three are fine, and nothing distinguishes them. That is the problem stated exactly.
So the next person starts from nothing. And somebody who tries an existing artifact, gets nothing out of it, and gives up has nowhere to say so, which means the person after them spends the same weeks finding out the same thing.
This is a record of what people extracted, from which model, how, and what happened when they or anyone else checked it. Failures included, because those are the ones that save somebody a month.
Several people will mean different things by the same word and build different vectors for it. When they do, those sit side by side here. Nothing picks one, and nothing is ranked by a score. Nobody has done it here yet: every submission in this corpus is by one author, so what this shows today is the mechanism rather than the disagreement. What this is and where it came from.