Accuracy evidence

What we compared against, and where it differed

Statory results compared against SPSS and Amos output on the same data. Every number below is read automatically from the comparison records in the repository — no one can edit them by hand.

320 values compared · 292 agree · 0 disagree

Summary

Analysis familyEnginesValues comparedAgreeConvention differsNot comparable
Named by the comparison run7939300
Difference148474010
Relationship5251933
Structural213724
Basic Analysis210802
Prediction26204
Basic Statistics13300
IBM SPSS Statistics33234206523
IBM SPSS Amos2868600

Record: docs/parity/spss_diff.md · docs/parity/amos_diff.md

Text & decision — what we compared against

SPSS and Amos do not produce these numbers, so instead of “do the values agree?” we record what each engine was compared against — an external reference, a same-library reproduction, an accuracy measured on a labelled public sample, or nothing comparable yet.

EngineVerdictAgainst what · how
ahpMatches an external referenceFor a perfectly consistent matrix (a_ij = w_i/w_j) the answer follows from the definition — weights = w, lambda-max = n, CI = 0, CR = 0 (Saaty's consistency definition). Group aggregation is computed separately with numpy as the element-wise geometric mean (AIJ) and compared.
collocationMatches an external referenceNLTK 3.x `BigramAssocMeasures.pmi` (Church & Hanks, 1990) — the two implementations are run side by side on the same tokens and the same window (4) and compared.
deaMatches an external referenceThe envelopment engine is solved again in its **dual multiplier form** and compared — CCR (1978), BCC (1984).
dominant_strategyMatches an external referenceDominance is a one-line definition, so it is worked out by hand and compared — standard textbook definition.
electreMatches an external referenceConcordance and discordance indices worked out by hand and compared — ELECTRE I, Roy (1968).
fuzzy_topsisMatches an external referenceChen, C.-T. (2000). "Extensions of the TOPSIS for group decision-making under fuzzy environment." *Fuzzy Sets and Systems*, 114(1), 1–9. — the formulation in that paper is set up and solved independently inside this file (the engine is not imported).
nash_equilibriumMatches an external referenceEquilibria solved independently with nashpy and compared — Nash (1950).
prometheeMatches an external referenceValues compared against pymcdm `PROMETHEE_II` (usual preference function) — Brans & Vincke (1985).
sawMatches an external referenceSimple additive weighting is a one-line definition, so it is worked out by hand and compared — MacCrimmon (1968).
sentiment_en_accuracyMeasured on a labelled sampleMeasured on a labelled sample - UCI Sentiment Labelled Sentences, 500 sentences (250 positive, 250 negative, CC BY 4.0) - accuracy 66.0%, positive F1 0.786, negative F1 0.678, 106 judged neutral. Kotzias (2015). The domain is product, movie and restaurant reviews, which differs from interview and survey text, and the sample carries no neutral label - read it as a reference figure.
sentiment_ko_accuracyMeasured on a labelled sampleAccuracy sample - NSMC, 500 labelled reviews (250 positive, 250 negative, CC0) - accuracy 48.4%, positive F1 0.714, negative F1 0.496, 216 judged neutral. The domain is movie reviews, which differs from interview and survey text - read it as a reference figure.
sequential_gameMatches an external referenceBackward induction worked out by hand and compared — subgame-perfect equilibrium, Selten (1965).
shapley_valueMatches an external referenceShapley values computed independently from the permutation-average definition and compared — Shapley (1953).
social_choiceMatches an external referenceBorda and Condorcet worked out by hand from their definitions and compared — Borda (1781), Condorcet (1785).
sociometryMatches an external referenceDegree, reciprocal choices and isolates counted independently with NetworkX and compared — Moreno (1934).
text_comprehensiveSame-library reproductionEach component measure (frequency, TF-IDF, sentiment) calls the same library again — a reproduction. There is no outside standard for the bundle as a whole.
topic_modelingSame-library reproductionThe engine and the reference both call scikit-learn `LatentDirichletAllocation` — this is a reproduction of the same library, not a comparison against an outside standard.
topsisMatches an external referenceValues compared against pymcdm `TOPSIS` (vector normalization) — Hwang & Yoon (1981).
vikorMatches an external referenceValues compared against pymcdm `VIKOR` (v = 0.5) — Opricovic & Tzeng (2004).
Total19Matches an external reference 15 · Same-library reproduction 2 · Measured on a labelled sample 2

Record: backend/tests/parity/family_reference.json

Values not counted as agreeing — and why

These values were not counted as agreeing. "Convention differs" means the values are supposed to differ (the tools compute them differently); "not comparable" means the reference tool does not report that value, or we could not extract it.

Show details (28)
  • Not comparabledescriptive · 중위수

    SPSS DESCRIPTIVES 는 중위수를 내지 않는다 — 대조 불가가 정상

  • Not comparablereliability · 표준화 α

    SPSS 는 /MODEL=ALPHA 만으로는 표준화 α 를 내지 않는다

  • Not comparablepaired_ttest · 자유도

    엔진 대응표본 표에 자유도 열이 없다(N 만 싣는다)

  • Not comparableone_sample_ttest · t

    대조 불가 · 동결 SPSS 출력 2본에 해당 표가 없다(Statory 산출은 정상)

  • Not comparableone_sample_ttest · 유의확률

    p 허용오차 바닥 0.0005(표시 3자리)

  • Not comparableone_sample_ttest · 자유도

    엔진 일표본 표에 자유도 열이 없다(N 만 싣는다)

  • Not comparablelogistic_regression · X_IND B

    대조 불가 · 동결 SPSS 출력 2본에 해당 표가 없다(Statory 산출은 정상)

  • Not comparablelogistic_regression · X_IND 유의확률

    p 허용오차 바닥 0.0005(표시 3자리)

  • Not comparablelogistic_regression · X_IND Exp(B)

    대조 불가 · 동결 SPSS 출력 2본에 해당 표가 없다(Statory 산출은 정상)

  • Convention differslogistic_regression · X_IND Wald

    SPSS 는 Wald χ²(=Z²), 엔진은 Wald Z 를 싣는다

  • Convention differsmoderation · 상호작용 B

    MR-6 ③ · FIX-4 ④ — SPSS 구문은 z 표준화 곱, 엔진은 평균중심화 곱(B_SPSS = B_엔진 × σ_x × σ_w)

  • Convention differsmoderation · 상호작용 베타

    MR-6 ③ — 엔진은 표준화 변수의 곱(z_x·z_w, Aiken & West), SPSS 구문은 곱항 자체를 표준화한다

  • Not comparablewilcoxon · 유의확률

    p 허용오차 바닥 0.0005(표시 3자리)

  • Not comparablewilcoxon · 효과크기 r

    SPSS 는 Wilcoxon 효과크기 r 을 내지 않는다

  • Not comparablefriedman · 유의확률

    p 허용오차 바닥 0.0005(표시 3자리)

  • Not comparablefriedman · Kendall W

    NPAR TESTS /FRIEDMAN 은 Kendall W 를 내지 않는다

  • Not comparablemcnemar · 불일치 셀 b

    SPSS 검정 통계량 표에는 불일치 셀 수가 없다(교차표에 있다)

  • Not comparablemann_whitney · 유의확률

    p 허용오차 바닥 0.0005(표시 3자리)

  • Not comparablecluster_analysis · X_IND F

    군집 번호는 SPSS 와 다를 수 있으나 변수별 F 는 번호와 무관하다

  • Not comparablecluster_analysis · X_MOD F

    군집 번호는 SPSS 와 다를 수 있으나 변수별 F 는 번호와 무관하다

  • Not comparablecluster_analysis · MEDIATOR F

    군집 번호는 SPSS 와 다를 수 있으나 변수별 F 는 번호와 무관하다

  • Not comparablecluster_analysis · 군집별 사례 수

    군집 번호가 SPSS 와 같다는 보장이 없어 번호로 맞대지 않는다

  • Convention differsdiscriminant · 집단 A 중심값

    판별함수 방향 — 여덟 칸이 통째로 반대 부호(크기는 일치). SPSS 부호 규칙은 출력에 비표준화 계수가 없어 미확인 — 판정 대기

  • Convention differsdiscriminant · 집단 B 중심값

    판별함수 방향 — 위와 같은 뒤집힘

  • Not comparablesurvival_km · 평균 생존시간

    표시 0자리 — 허용오차를 반올림 폭 0.5 로 넓힘

  • Not comparablesurvival_km · 중위 생존시간

    대조 불가 · 양쪽 미확보 — SPSS 출력에도 없고 엔진도 내지 않는다. 표준 보고 항목이므로 Statory 산출 보강 후보

  • Not comparablesurvival_cox · X_IND B

    대조 불가 · 엔진이 계수 B 를 내보내지 않는다(계산은 하고 위험비 HR 로만 싣는다) — 엔진 보강 대기

  • Not comparablesurvival_cox · 전체 카이제곱

    대조 불가 · 양쪽 미확보 — SPSS 출력에도 없고 엔진도 내지 않는다. 표준 보고 항목이므로 Statory 산출 보강 후보

By reference tool

IBM SPSS Statistics234 values compared · 206 agree · 0 disagree
ancova5500
cluster_analysis5104
correlation1100
crosstab3300
descriptive8701
discriminant8620
efa_paf_varimax393900
efa_pc_varimax414100
friedman4202
hierarchical_regression4400
independent_ttest3300
kruskal_wallis4400
logistic_regression4013
mann_whitney3201
manova8800
mcnemar3201
mediation_step2_a2200
mediation_step3_b_cprime3300
mixed_anova212100
moderation3120
multiple_regression131300
one_sample_ttest3003
one_way_anova8800
one_way_anova(agri)5500
ordinal_regression2200
paired_ttest4301
reliability2101
repeated_measures_anova7700
spearman_correlation1100
survival_cox4202
survival_km2002
two_way_anova8800
wilcoxon3102
IBM SPSS Amos86 values compared · 86 agree · 0 disagree
measurement414100
structural454500

Engine grades

Grades are derived automatically from golden tests and invariance measurements. No one assigns them by hand.

Verified 114Conditional 16Unverified 0

How each engine was verified

Every engine has at least one form of evidence. Counts by kind:

Compared against a reference implementation148
Checked by hand calculation55
Golden test exists, reference not attributable22
Compared against a textbook example16
Compared against IBM SPSS Statistics output7
Compared against IBM SPSS Amos output2

How to cite Statory

Copy the line below into your paper or report.

Statory. (2026). Statory (Version 1.5.0) [Statistical analysis software]. https://www.statory.org

Reproducibility — will the same result come back

  • Each analysis records which engine ran it and with which options, alongside the result.
  • Norms are stored under a version name (e.g. 20260906), so you can trace which norm a past result used.
  • The comparison records themselves are frozen in the repository and re-run on every release — if a value changes, the release is blocked.

SPSS and Amos are trademarks of their respective owners. Statory is not affiliated with, sponsored by, or endorsed by them. This page reports the output each tool produced from the same data, side by side.

Back to the guide