The Unjournal (Unjournal.org)

@unjournal.bsky.social

Researchers, practitioners, & open science advocates building a better system for research evaluation. Nonprofit. We commission public evaluation & rating of hosted work. To make rigorous research more impactful, & impactful research more rigorous.

Starts Friday: the Animal Futures Tournament on Metaculus. Forecast 16 decision-relevant questions on farmed animal welfare, alternative proteins, wild animal welfare, AI, and animal futures. The Unjournal is partnering with Metaculus and Sentient Futures through our Pivotal… http://dlvr.it/TSzF8b

Animal Futures Tournament — The Unjournal

The Animal Futures Tournament is on Metaculus. Forecast decision-relevant questions on farmed animal welfare, alternative proteins, wild animal welfare, AI, and animal futures.

dlvr.it

Are IO/scanner-data price-elasticity estimates reliable? "Lalonde-style" evidence raises doubts: observational scanner variation doesn't reproduce experimental elasticities (Bray, Sanders, Stamatopoulos). Unjournal evals: https://unjournal.pubpub.org/pub/evalsumbraybray/ #EconSky

Scale use adjustments in wellbeing measures; Benjamin, Kimball; Kaiser response. Unjournal workshop

See the workshop summary and interactive transcript and resources here : https://uj-wellbeing-workshop.netlify.app/transcript 0:00 Intro 0:08 Dan Benjamin begins: paper overview and motivation 1:21 Scale-use heterogeneity: what problem is the paper solving? 3:39 Why calibration questions are needed 6:30 Evidence that scale-use heterogeneity is real 8:00 Panel evidence: scale use changes over time 12:05 Height and weight as objective validation cases 12:42 Regression results: correcting coefficients 13:12 Why the correction matters for policy-relevant ratios 15:12 Method of moments estimator 16:12 Validation: better alignment between subjective and objective measures 18:12 Practical recommendation: add calibration questions to surveys 23:12 Robustness and comparison of calibration approaches 28:12 Implications for pricing exercises and evaluation 33:12 Caspar Kaiser’s response: strengths of the paper 34:53 Response consistency vs common perception trade-off 35:41 CHOPIT vs method of moments 38:41 Benjamin on distributional assumptions and identification 40:41 Practical implications for charity evaluation 42:41 Adding calibration questions in practice 45:41 Generalizability to LMICs 47:11 Outro Here: Dan Benjamin presents joint work on scale-use heterogeneity in subjective wellbeing measurement: the problem that two respondents may mean different things when they give the same life-satisfaction score. The paper’s core proposal is to use calibration questions or vignettes to estimate how people translate their internal states onto survey scales, so analysts can correct for differences in both the level and spread of scale use. Using new data from the Understanding America Study, Benjamin offers evidence that scale-use heterogeneity is real, that it changes over time, and that correcting for it can materially change some policy-relevant estimates. Benjamin also emphasizes a practical recommendation for implementers: adding a small number of calibration questions to surveys may be a tractable way to improve subjective wellbeing measurement in real-world evaluation settings. Caspar Kaiser’s response is positive but demanding. He welcomes the stronger theoretical foundations and especially the objective-subjective validation, but presses on several important issues. First, he notes that while many coefficients do not move dramatically, the ratios between them can be fragile, and those ratios are exactly what many workshop participants care about for evaluation and prioritization. Second, he highlights a tradeoff between “response consistency” and “common perception” assumptions: it may be plausible that people use the same scale for self-reports and vignettes, but less clear that they interpret vignettes in the same way. Third, he asks for clearer positioning relative to CHOPIT-style approaches, more sensitivity analysis, and practical tooling such as usable software packages. The talk suggests that scale-use correction could be highly important when using wellbeing data to compare interventions or estimate money-equivalent tradeoffs. But Kaiser also stresses that most of the evidence here is from US data, while many of the highest-stakes practical applications are in low- and middle-income countries. So the session ends with both a methodological advance and an (feasible?) agenda: if wellbeing measures are going to guide real intervention choices in LMICs, we might benefit from calibration data, implementation tools, and direct evidence from those settings rather than assuming the corrections generalize automatically. https://uj-wellbeing-workshop.netlify.app/about unjournal.org

youtube.com

Unjournal Wellbeing Measurement Workshop: Scale Use Discussion

Can we really compare interventions using self-reported wellbeing? This discussion looks at a basic but important problem: people may use wellbeing scales differently, and interventions themselves may even change how people interpret and answer those questions. If so, a “1-point gain” may not mean the same thing across programs or groups. Caspar Kaiser, Matt Lerner, Peter Hickman, Miles Kimball, and others discuss what this means for comparing psychotherapy, cash transfers, health interventions, and other programs. The conversation also asks whether alternatives like DALYs or cost-benefit analysis avoid the same problem, or just hide it differently. The Unjournal: https://unjournal.org 0:00 Is scale use relevant to choosing between interventions? 0:37 Caspar: psychotherapy vs cash transfers may affect scale use differently 1:17 How interventions might change interpretation of scales 2:03 We still do not know intervention-specific effects well 2:30 Similar concerns may apply to other metrics too 3:15 Counterfactuals to wellbeing measures 3:39 Revealed preference as an alternative 4:08 Dollar-based cost-benefit analysis and its assumptions 4:42 Evaluate wellbeing measures relative to the next-best alternative 5:16 Ad hoc revealed-preference fixes might carry over 5:54 Matt: cash transfers as the benchmark 6:26 Compare to cash using wellbeing data directly 7:11 Matt: ordinality is the main concern 7:53 When preserving rankings may be enough 8:33 When cardinal information still matters 8:57 Peter: DALY values can change program rankings 9:22 What scale correction would mean for Charity Governance 9:47 Triangulating across methods rather than relying on one formula

youtube.com

Highlights/Samples from Unjournal.org's "Wellbeing Measures for LMIC Interventions" workshop

See https://uj-wellbeing-workshop.netlify.app/about.html and https://uj-wellbeing-workshop.netlify.app/transcript for more detailed discussion and resources from the workshop Topics covered include: • why scale-use heterogeneity matters for subjective well-being measures • whether linear life-satisfaction scales can be trusted across people and contexts • how WELLBYs compare to DALYs, stated preferences, and revealed-preference approaches • practical problems in converting DALYs to WELLBYs • what evaluators and funders actually need from well-being metrics • belief elicitation and open questions for future research Chapters: 0:00 Opening and workshop goals 7:08 Peter Hickman: stakeholder perspective from Charity Governance 16:34 Matt Lerner: stakeholder perspective from Founders Pledge 44:27 Caspar Kaiser: key measurement concerns 1:02:22 Scale use discussion 1:12:27 DALY/WELLBY history and background 1:23:19 Samuel Dupret and Julian Jamison: DALY-WELLBY conversion 2:08:02 Dan Benjamin and Miles Kimball: research presentation 3:07:27 Evaluator responses and critique 3:28:30 Belief elicitation and pivotal questions 3:47:27 Practitioner panel and open discussion This recording is part of The Unjournal’s broader work on high-stakes empirical and conceptual questions in research evaluation, philanthropy, and global priorities. The Unjournal: https://unjournal.org

youtube.com

For journal-independent evaluation to matter for careers and trust, ratings and reviews need to be visible where researchers and institutions actually look—Google Scholar, repositories, indexing services. That's why we analyzed where Unjournal-evaluated research is hosted. 🧵

Honorable Mention Lottery: Drawing Tuesday, February 10 Two of five honorable mention evaluators will win an extra $500 each. Why a lottery? Real incentives, not token recognition — larger payments to fewer people reduces transaction costs for everyone. Results announced… http://dlvr.it/TQr5fQ

Honorable Mention Lottery — The Unjournal 2024–25 Evaluator Prize

Transparent random draw to select two $500 prize winners from the five honorable mentions in The Unjournal's 2024–25 Evaluator Prize.

dlvr.it

🎥 NEW VIDEO: The Unjournal Process — How We Evaluate Research Watch our 6-step process for rigorous, transparent peer review: • Expert evaluators compensated $350-450/review • 3-5 week turnaround • Public evaluations with DOIs • Author response before publication… http://dlvr.it/TQqbpC

- YouTube

Enjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.

youtube.com

Designing & building a new interface for Unjournal research evaluations. Would love your thoughts & feedback. A very preliminary version:

Evaluation Form - Unjournal

Please aim to write a report up to the standards of a high-quality referee report for a traditional journal. Consider standard guidelines as well as The Unjournal's emphases. Remember to address any specific considerations mentioned by the evaluation manager, including in our bespoke evaluation notes. Please provide a concise summary of your evaluation below. Otherwise, please write your evaluation here, provide a link to it, or let us know you have emailed it. If you are linking or sending a file, we prefer the 'native format' (word, latex, markdown, bibtex, etc.) and not a pdf (but if necessary we can handle a pdf). These are percentile ratings relative to a reference group: serious research in the same area that you have encountered in the last few years. A rating of 50 means the paper is at the median of this reference group; 80 means it is in the top 20%; 20 means only 20% of comparable work is worse. See the guidelines on quantitative metrics for details. Tip: Use the Calibrate button above to practice rating sample papers and check your calibration. Guidance Judge the quality of the research heuristically. Consider all aspects of quality, credibility, importance to future impactful applied research, and practical relevance and usefulness, importance to knowledge production, and importance to practice. Benchmark: serious research in the same area encountered in the last three years. Guidance Consider the following: Are the methods well-justified and explained? Are they a reasonable approach to answering the question(s) in this context? Are the underlying assumptions reasonable? Are the results and methods likely to be robust to reasonable changes in assumptions? Does the author demonstrate robustness? Did the authors take steps to reduce bias from opportunistic reporting and questionable research practices? Guidance To what extent does the project contribute to the field or to practice, particularly in ways relevant to global priorities and impactful interventions? Focus on "improvements that are actually helpful" (applied stream) Originality and cleverness should be weighted less than typical journals — we focus on impact More weight on "contribution to global priorities" than "contribution to academic field" Do the paper's insights inform beliefs about important parameters and intervention effectiveness? Does the project add useful value to other impactful research? Sound, well-presented null results can also be valuable Guidance Are goals and questions clearly expressed? Are concepts clearly defined and referenced? Is the reasoning transparent? Assumptions explicit? Are all logical steps clear and correct? Does the writing make arguments easy to follow? Are conclusions consistent with the evidence presented? Do authors accurately characterize evidence and its support for main claims? Are data and analysis relevant to the arguments? Are tables, graphs, diagrams easy to understand (no major labeling errors)? Guidance This covers several considerations: Replicability, reproducibility, data integrity: Would another researcher be able to perform the same analysis and get the same results? Are methods explained clearly enough for credible replication? Is code provided? Is data source clear and as available as reasonably possible? Consistency: Do numbers in the paper and code output make sense? Are they internally consistent throughout? Useful building blocks: Do authors provide tools, resources, data, and outputs that might enable future work and meta-analysis? Reference: COS TOP Guidelines — a framework for evaluating transparency across 8 dimensions including data, code, materials, and preregistration. Guidance Is the topic and approach useful to global priorities, cause prioritization, and high-impact interventions? Does the paper consider real-world relevance, policy, and implementation questions? Are the setup, assumptions, and focus realistic? Do authors report results relevant to practitioners? Do they provide useful quantified estimates (costs, benefits) for impact quantification? Do they communicate in ways policymakers can understand without misleading oversimplification? Guidance Where should this paper be published based on merit alone? Imagine a journal process that is fair, unbiased, and free of noise — where status, connections, and lobbying don't matter. 0: Won't publish / little to no value 1: OK / Somewhat valuable journal 2: Marginal B-journal / Decent field journal 3: Top B-journal / Strong field journal 4: Marginal A-journal / Top field journal 5: A-journal / Top journal Non-integer scores encouraged (e.g., 4.6, 2.2). 0 Won't publish 1 OK 2 Marginal B 3 Top B 4 Marginal A 5 Top A Guidance Where will this research actually be published? If already published and you know where, report the prediction you would have given absent that knowledge. 0: Won't publish / little to no value 1: OK / Somewhat valuable journal 2: Marginal B-journal / Decent field journal 3: Top B-journal / Strong field journal 4: Marginal A-journal / Top field journal 5: A-journal / Top journal Non-integer scores encouraged (e.g., 4.6, 2.2). 0 Won't publish 1 OK 2 Marginal B 3 Top B 4 Marginal A 5 Top A This section is meant to help practitioners use this research to inform their funding, policymaking, and other decisions. It is not intended as a metric to judge the research quality per se. This is mainly relevant for empirical research. If 'claim assessment' does not make sense for this paper, please consult the evaluation manager, or skip this section. I. Identify the most important and impactful factual claim this research makes — e.g., a binary claim or a point estimate or prediction. III. [Optional] What additional information, evidence, replication, or robustness check would make you substantially more (or less) confident in this claim? IV. [Optional] Identify the important implication of the above claim for funding and policy choices. To what extent do you believe this implication? How should it inform policy choices? + Add Another Claim We generally incorporate this into the 'abstract' of your evaluation (see examples at unjournal.pubpub.org). Your comments here will not be public or seen by authors. Please use this section only for comments that are personal/sensitive in nature. Please place most of your evaluation in the public section. Please disclose your use of AI/LLM tools in this evaluation. AI tools may be used for specific tasks (literature search, methodology checks, writing assistance) but not for generating overall evaluations or ratings. You must independently verify any AI-generated suggestions. See recommended tools → Responses to these will be public unless you mention in your response that you want us to keep them private. Feedback (responses below will not be public or seen by authors) Would you be interested in discussing this with other evaluators and writing a joint report? See bit.ly/UJevalcollab. It will come with some additional compensation. If you and other evaluators are interested, we may follow up to arrange an (anonymous) discussion space or synchronous meetings. Tick here if you would like to be part of our 'evaluator pool' in the future. To be contacted for compensated evaluation work when research comes up in your area. To expedite this, fill out the EOI form at Join the Unjournal. Do you have any other suggestions or questions about this process or The Unjournal? (We will try to respond, and incorporate your suggestions.)

dlvr.it