{
  "slug": "dlaczego-rankingi-modeli-sie-roznia",
  "ranking_id": "D01b",
  "edition": "1.0",
  "date": "2026-09-26",
  "series": {
    "pl": "Lider i jego warsztat",
    "en": "The leader and their workspace"
  },
  "number": 5,
  "title": {
    "pl": "Dlaczego modele tworzą różne rankingi tych samych rzeczy",
    "en": "Why models rank the same things differently"
  },
  "intro": {
    "pl": "Te same oceny, inny zwycięzca. Sprawdzony przykład pokazuje wpływ wag i wzoru oraz granice porównywania modeli.",
    "en": "The same ratings, a different winner. A checked example shows how weights and formulas affect rankings, and where model comparisons fall short."
  },
  "source_dates": [
    "2026-08-05",
    "2026-08-06",
    "2008"
  ],
  "checked_at": "2026-09-26",
  "sources": [
    {
      "title": "Historyczna analiza rozbieżności i raport porównawczy",
      "access": "internal",
      "date": "2026-08-06",
      "scope": "Both full comparison documents; numerical fields and ranking table used for diagnostic arithmetic; private rankings not published"
    },
    {
      "title": "OECD/JRC Handbook on Constructing Composite Indicators, sections 1.6–1.7",
      "url": "https://www.oecd.org/en/publications/handbook-on-constructing-composite-indicators-methodology-and-user-guide_9789264043466-en.html",
      "accessed_at": "2026-09-26",
      "date": "2008"
    }
  ],
  "evidence_scope": {
    "pl": "Fikcyjny przykład i sprawdzona arytmetyka. Bez nowego benchmarku modeli i bez pełnego audytu historycznych tez naukowych.",
    "en": "Fictional example and checked arithmetic. No new model benchmark or full audit of historical scientific claims."
  },
  "update_trigger": {
    "pl": "Błąd rachunku, nowe dane źródłowe lub zmiana metody.",
    "en": "An arithmetic error, new source data or a method change."
  },
  "ai": {
    "authoring": "Codex",
    "source_selection": "Claude Code + Codex",
    "human_review": "post-publication withdrawal by Lech R. Rustecki",
    "independent_model_review": false
  },
  "interactive": "rankings-diagram",
  "schema_version": 1,
  "language": "en",
  "url": "https://www.l00p.ai/en/resources/series/dlaczego-rankingi-modeli-sie-roznia/",
  "markdown_sha256": "6fda383fbc4e75f2483977c3ad1ec947f9d5274f0ee513b6d06d4aa660d4bf75",
  "sections": [
    {
      "id": "rozdzial-01",
      "title": "01 / Two lists may answer two questions",
      "url": "https://www.l00p.ai/en/resources/series/dlaczego-rankingi-modeli-sie-roznia/#rozdzial-01",
      "markdown": "You ask two models for the best projects. One chooses breakthrough ideas; the other chooses projects that can be completed this year. Both provide an ordered list and persuasive reasoning. Before deciding which model “knows more”, check what each understood by best.\n\nOur starting point is two internal comparison documents dated 6 August 2026 and stored assessments from the preceding day. They concerned research priorities. We preserve the methodological lesson without publishing private rankings. This is neither a current model-quality test nor an assessment of today's scientific problem status. Model names recorded in the archive are not treated here as verified service specifications."
    },
    {
      "id": "rozdzial-02",
      "title": "02 / A ranking comes from several decisions",
      "url": "https://www.l00p.ai/en/resources/series/dlaczego-rankingi-modeli-sie-roznia/#rozdzial-02",
      "markdown": "The result depends on candidates, criterion definitions, assessments, scales, weights and the way numbers are combined. Sources and their verification dates also matter. Two lists with similar headings may differ at every one of these levels.\n\nStart with a shared vocabulary: do two names refer to the same object and scope? Then establish whether candidates were assessed using the same instructions. “Feasible” might mean a small team's work for a week or the prospect of progress over a decade. A shared zero-to-ten scale does not repair that difference.\n\nIf you want to isolate the effect of a formula, keep the assessments and candidates unchanged. If you change everything at once, you see different outcomes but cannot tell what caused them."
    },
    {
      "id": "rozdzial-03",
      "title": "03 / A neutral example: a two-tenths lead",
      "url": "https://www.l00p.ai/en/resources/series/dlaczego-rankingi-modeli-sie-roznia/#rozdzial-03",
      "markdown": "Take two fictional projects. A scores 10 for value, 10 for scientific effect and 1 for feasibility. B scores 8, 7 and 6 respectively. These are illustrative ratings, not measurements or probabilities. For the exercise, we treat point differences as comparable.\n\nThe first rule assigns 40% to value, 30% to scientific effect and 30% to feasibility. The weights sum to 100%:\n\n- A: 0.40 × 10 + 0.30 × 10 + 0.30 × 1 = 7.3.\n- B: 0.40 × 8 + 0.30 × 7 + 0.30 × 6 = 7.1.\n\nA leads by 0.2 points. The weighted sum lets high assessments on two dimensions compensate for low feasibility. This is a property of the chosen rule. It does not prove that project A can be completed."
    },
    {
      "id": "rozdzial-04",
      "title": "04 / Change only the weights",
      "url": "https://www.l00p.ai/en/resources/series/dlaczego-rankingi-modeli-sie-roznia/#rozdzial-04",
      "markdown": "Move 10 percentage points from value to feasibility. The new weights are 30%, 30% and 40%; project assessments remain identical. A scores 6.4 and B scores 6.9. B now leads by 0.5 points. No new fact about the projects has appeared. The ranking designer's priorities have changed.\n\n![Same ratings, three rules. Weights 40/30/30: A 7.3, B 7.1. Weights 30/30/40: A 6.4, B 6.9. Mean of two ratings times feasibility/10: A 1.0, B 4.5. All data fictional.](https://www.l00p.ai/wydawnictwo/dlaczego-rankingi-modeli-sie-roznia/reguly-en-v1.svg)\n\nOriginal numerical diagram, Codex / L00P.AI. Shared 0–10 scale; bar length represents the score. Compare order within each rule. This is not a model benchmark.\n\nThe diagram also shows a second experiment, explained in the next section. Every bar starts at zero and the full scale is 10 points. [Download the data and formulas as CSV](/wydawnictwo/dlaczego-rankingi-modeli-sie-roznia/przyklad-v1.csv). The file contains six results: two projects under three rules.\n\nDo not choose weights only after seeing the winner. Agree them with the decision owner, then show whether reasonable alternatives change the order. If they do, the result calls for a discussion about priorities, not another decimal place."
    },
    {
      "id": "rozdzial-05",
      "title": "05 / Multiplication changes the role of a low rating",
      "url": "https://www.l00p.ai/en/resources/series/dlaczego-rankingi-modeli-sie-roznia/#rozdzial-05",
      "markdown": "In a separate, simplified variant, average value and scientific effect, then multiply by feasibility divided by ten:\n\n- A: ((10 + 10) / 2) × (1 / 10) = 1.0.\n- B: ((8 + 7) / 2) × (6 / 10) = 4.5.\n\nThis is a new rule, not merely a change to the previous sum's weights. Low feasibility constrains the whole score; zero feasibility would make it zero. This simplified formula illustrates the mechanism and does not reproduce every component of the historical ranking. Compare the order within each rule; do not interpret point differences between rules as a measured loss of project value.\n\nThe [OECD/JRC handbook](https://www.oecd.org/en/publications/handbook-on-constructing-composite-indicators-methodology-and-user-guide_9789264043466-en.html), sections 1.6–1.7, recommends documenting the rationale for weighting, aggregation and sensitivity analysis. It also discusses compensating for a weakness in one dimension with a strength in another. If a condition is mandatory, apply an [eligibility criterion](/en/resources/series/jak-wybierac-system-do-dzialajacej-organizacji/) first, rather than expecting the formula alone to enforce it."
    },
    {
      "id": "rozdzial-06",
      "title": "06 / What the historical comparison needs to correct",
      "url": "https://www.l00p.ai/en/resources/series/dlaczego-rankingi-modeli-sie-roznia/#rozdzial-06",
      "markdown": "The source analysis usefully explained the difference between adding and multiplying ratings, but attributed divergence too firmly to the formula alone. The data we read also contain different assessments and candidate sets. Without a shared basis, we cannot quantify how much of the change in order came from the method.\n\nRecalculating the stored components did not exactly reproduce every stored result. Without further data, we cannot establish whether rounding, different inputs or an error caused this. The archive remains unchanged. The public example has its own complete inputs and checked arithmetic.\n\nThe second correction concerns the name “ROI”. Calling a point index return on investment does not make it a financial return measure. A rating of 1/10 is not automatically a 10% probability of success. That interpretation would require justified calibration, which is absent here. Nor do we repeat an author's praise of their own comparison as an independent review."
    },
    {
      "id": "rozdzial-07",
      "title": "07 / Ask for inputs, not just a verdict",
      "url": "https://www.l00p.ai/en/resources/series/dlaczego-rankingi-modeli-sie-roznia/#rozdzial-07",
      "markdown": "A useful model task should specify a common list, criterion meanings, source date, missing-data treatment and assessment format. Mark an unknown assessment as missing with a reason; do not replace it with zero. Separate assessment from calculation: a simple script can apply the same formula to everyone.\n\nCheck separately: different weights on the same data, different assessments under the same formula, and the effect of adding or removing candidates. Show ties explicitly. Record which conclusions remain stable and which depend on assumptions. Correct arithmetic does not validate the assessments.\n\nOne decision and two alternatives are enough to start a conversation. We can help establish a common basis for comparison while keeping project data private. The goal is an explainable decision, not a table in which a favourite product always wins."
    },
    {
      "id": "rozdzial-08",
      "title": "08 / Sources, numbers and AI involvement",
      "url": "https://www.l00p.ai/en/resources/series/dlaczego-rankingi-modeli-sie-roznia/#rozdzial-08",
      "markdown": "We read both historical comparison documents, the assessment table and the machine-readable dataset's numerical fields. Verification on 26 September 2026 covers arithmetic and method, not a complete new audit of historical scientific claims. The OECD/JRC handbook dates from 2008; we checked the cited sections and do not present it as a new publication.\n\nCodex prepared the PL/EN text, CSV and deterministic SVG. All six example results were recalculated. No third-party illustrations, new model benchmark or independent second-model review were used. A human can withdraw the material after publication. An arithmetic error, new source data or a method change requires a new acceptance review and an explicit edition."
    }
  ],
  "media": [
    {
      "url": "https://www.l00p.ai/wydawnictwo/dlaczego-rankingi-modeli-sie-roznia/reguly-en-v1.svg",
      "kind": "diagram_in_code",
      "depicts_real_measurement": false
    }
  ]
}
