Systemic Altruism
An inspectable method · version 0.1

Potential is a question.
Evidence is the test.

We ask how an organization could change the conditions that produce a problem. We then show what is documented, what is a judgment, and what still needs review.

All initial scores are AI-assisted preliminary editorial design hypotheses, without independent human validation yet. They are not causal impact estimates, cost-effectiveness estimates, verified outcomes, or a certification of any organization.

What systemic change means here

A mechanism changes rules, incentives, institutions, shared knowledge or decision-making power in ways that could persist beyond repeated delivery of a service. Direct service may save lives, build trust, generate evidence or enable later structural change. A lower structural score says nothing about the moral importance of that work.

The six dimensions

Structural depth · default weight 30%

Does the mechanism change rules, incentives, power or institutions? See each profile for its specific rationale and evidence gaps.

Institutional reach · default weight 20%

How widely can the documented mechanism influence a system? See each profile for its specific rationale and evidence gaps.

Durability · default weight 20%

Could the change persist without continued delivery by the charity? See each profile for its specific rationale and evidence gaps.

Scalability · default weight 15%

Can the mechanism spread beyond its initial setting? See each profile for its specific rationale and evidence gaps.

Funding additionality · default weight 10%

Would another donation enable work that otherwise would not happen? See each profile for its specific rationale and evidence gaps.

Community agency · default weight 5%

Do affected communities have documented decision-making power? See each profile for its specific rationale and evidence gaps.

The 0–4 rubric

  • 0: reviewed evidence does not identify this mechanism or capacity in the assessed work.
  • 1: limited or local pathway; mostly ongoing delivery.
  • 2: a plausible, bounded pathway with some institutional or replication potential.
  • 3: an explicit pathway to broader, persistent change.
  • 4: structural change is central to the documented mechanism, with a broad or durable design.
  • Unreviewed: insufficient review to assign a value. It is never silently converted to zero.

The rubric is a qualitative mechanism assessment. Higher ambition is not evidence of success. Organizational self-reports can describe a mechanism without proving attribution, additionality or outcomes.

How the range and order work

Normalize your non-negative weights to 100%. For each reviewed axis, multiply its 0–4 value by its weight and divide by 4. Their sum is the lower bound. The upper bound adds the full weight of all unreviewed axes. Weighted coverage is the sum of the reviewed weights. An all-unreviewed profile spans 0–100; it is not rated zero.

Default ordering uses the documented lower bound with alphabetical tie-breaking. This is a conservative display order within the starting cohort, not a total order of effectiveness. Wide or overlapping ranges mean the available review cannot separate organizations. Changing weights reflects values, not new evidence. These envelopes only capture missing-axis uncertainty: scored axes also contain qualitative and factual uncertainty.

Selection and coverage

The starting cohort contains 40 organizations: four GiveWell top-charity programs, five Giving Green climate recommendations, ten Animal Charity Evaluators recommendations, three historical program reviews, and 18 illustrative organizations across public health, poverty, governance, rights and knowledge. Evaluator inclusion must be checked against the linked dated source; it is not endorsement of our method. Large humanitarian charities and many locally led organizations remain outside this cohort. We do not claim global coverage.

GiveWell evaluates specified programs, so a program's evidence cannot automatically be generalized to the entire organization. Giving Green already explicitly studies systemic change, and ACE examines policy and industry mechanisms. Our contribution is an open cross-cause interface with visible judgments and adjustable priorities.

Evidence quality is separate

Profiles distinguish evaluator-supported program evidence from organizational descriptions. Neither label is a causal-confidence score. Before publishing any stronger outcome claim we need independent evaluations, counterfactual analysis, attribution, cost and funding-gap evidence, and affected-community review. Additionality and agency remain unreviewed where those audits are absent.

Review and corrections

Send a public source and precise correction through the evidence issue form. Maintainers review proposed changes in a pull request with a source, date, scope and rationale for each affected axis. Model-assisted extraction can prepare a local review draft; no model output can directly change published rankings. All changes remain visible in version history.

Conflicts and governance

The tool is a Systemic Altruism project. Its maintainers' movement affiliations create a potential selection bias. Affiliated organizations are not scored in this initial cohort. No payments, donation commissions or sponsorships determine inclusion or scores. Future affiliations and funding must be disclosed before release. Independent and community review are still needed.