PanelSimSign inCreate account
Case study · ICLR · the bundled sample paper

From likely reject to likely accept in one revision.

Contrastive Calibration for Retrieval-Augmented Summarization, a 1,974-word draft submitted to PanelSim as an Overleaf export, targeted at International Conference on Learning Representations. Every number below is the output of PanelSim on that project: the first report, the revision run, and the second report on the revised project. The venue and the paper are the ones bundled with the product, so you can open both reports yourself.

Read as Natural language processing, faithfulness of retrieval-augmented summarization, an empirical paper. The panel drew 10 reviewers chosen for it, each scoring every criterion, from senior researcher in faithfulness of abstractive summarization to researcher in calibration of language models, and a senior area chair in retrieval-augmented generation.

First report
1%
Likely reject, mean 4.96 · AC: Reject
Edits executed
11 of 12
78 lines added or changed, 4 files generated
Author time
~10 h
6 measurements and rewrites only the author could make
Second report
79%
Likely accept, mean 6.39 · AC: Accept (poster)

Where the panels landed, before and after

accept threshold 6before 4.96after 6.3912345678910more0
first report, 2000 simulated panelssecond report, revised paper
1

Day 0: the first report

The draft was uploaded as it stood: three results tables with single point estimates, two figures, eighteen references, no limitations section, no reproducibility statement. PanelSim flagged 10 mechanical findings before any review was written, then simulated ten reviewers across six cohorts.

The panel put the paper at 1% acceptance with a predicted mean of 4.96 against a threshold of 6: Likely reject. The reviews agreed on the substance and disagreed on the score, which is what a real panel does.

21 distinct concerns after merging across reviewers, turned into 20 ranked edits. Open the first report.

RejectArea chair, confidence 4 of 5

The reviews converge on three issues that bear directly on the main claim: every number is a single point estimate (R1, R2), nothing attributes the gain to the contrastive term rather than to the extra training signal (R1, R3), and the faithfulness metric is shared with the post-hoc filtering baseline, which is therefore optimizing the metric directly (R1, R5). R3 finds the objective well motivated and R10 finds the paper readable; neither disputes the missing evidence. R7 holds the lowest score and I weigh that review over the enthusiasm of R10 because its concerns point to specific tables. In its current form the evidence does not support the claims at the level the venue expects, and I recommend rejection with a clear path to resubmission.

What the reviewers raised most

ConcernRaised bySeverity
Single point estimates throughout
Experiments
3/10Serious
No ablation of the negative construction
Method
3/10Serious
Faithfulness metric shares the entailment model with a baseline
Experiments
3/10Serious
Claims exceed the evidence
Introduction
3/10Serious
No limitations discussion
Whole document
3/10Serious
The perturbed summaries are assumed to be unsupported
Method
1/10Serious
Training cost of the contrastive term is not reported
Experiments
1/10Serious
Delta over sequence-level contrastive training is not isolated
Related Work
1/10Serious
2

Day 1: the plan, and the revision run

The author selected 12 of the ranked edits, 18.8 estimated author hours and 16 compute hours in the plan, and handed them back. PanelSim wrote the passages from the project's own content, replaced the sentences that overclaimed, computed the variance table from the recorded seeds in data/results.csv, produced the ablation figure and table with the measured endpoints and the two still-to-run variants marked as such, and verified the two references it proposed against Crossref before writing them into refs.bib. Predicted mean after the run: 5.92 (42%), before any of the author's own measurements.

#EditKindHoursResult
1
Add a limitations section
Written from the project content for: Add a limitations section.
text1 applied
main.tex
2
Scope the claims to what the tables show
Replaced the passage in place and added a scoped statement.
text1.5 applied
sections/intro.tex
3
State seeds and repeated runs
Written from the project content for: State seeds and repeated runs.
text0.5 applied
sections/experiments.tex
4
Cite or remove the unused entries
2 reference(s) verified against an open index, 0 discarded as unverifiable.
citation0.5 applied
sections/related.tex
5
Write a standalone caption for the overview figure
Replaced the passage in place and left the rest of the text untouched.
text0.25 applied
sections/intro.tex
6
Add a reproducibility statement
Written from the project content for: Add a reproducibility statement.
text1 applied
sections/experiments.tex
7
Describe the span tagger
Written from the project content for: Describe the span tagger.
text1 applied
sections/method.tex
8
Report variance and a paired test over five seeds
Recomputes Table 1 from results.csv over the recorded seeds with standard deviations and a paired test on the combined criterion.
experiment4 applied
sections/experiments.tex
panelsim_table_variance.tex
9
Discuss the closest contrastive faithfulness work
The references this edit calls for are already in the bibliography (welleck2020neural, cao2021cliff), added by an earlier edit or present in the project.
citation1.5 skipped
10
Add an ethics and impact statement
Written from the project content for: Add an ethics and impact statement.
text0.75 applied
main.tex
11
State the contribution one way throughout
Replaced the passage in place and added a scoped statement.
text0.75 applied
sections/conclusion.tex
12
Ablate the source of negatives
Ablation figure and table over the source of negatives, with the measured endpoints taken from results.csv and the intermediate rows marked as runs still to be made.
experiment6 applied
sections/experiments.tex
panelsim_fig_ablation.pdf, panelsim_fig_ablation.png, panelsim_table_ablation.tex

The revised project came back as a zip with a unified patch across 6 files, ready to upload to Overleaf or apply with patch -p1.

3

Days 2 to 4: what only the author could do

The report is explicit about the line between what PanelSim can produce and what needs the author's machines. The generated ablation script marked two rows as placeholders; the independent-judge table needed a second entailment model run; training cost and a larger passage set needed the GPU. The author did those, in about 10 hours, and folded the numbers into the revised project.

Follow-throughHours
Ran the two intermediate ablation variants (random negatives, in-document negatives) that the generated script had marked as placeholders and put the measured values into the ablation table.3
Scored every system with a second entailment model, added the column to results.csv and re-ran the judge table, then measured how often perturbed spans were supported by another passage (3.6 percent on 400 summaries).2.5
Measured training cost against the baseline on the same hardware and added the paragraph: 1.8 times the wall clock per step, 1.3 times the memory.1
Re-ran the main experiment with k = 20 and added the result.1.5
Restated the gradient locality claim as a statement about the loss under teacher forcing, as the theory reviewer asked.0.25
Cited the two bibliography entries that were never discussed where they belong.0.25
4

Day 5: the second report

The revised project, now 3,220 words, went through the same panel. 0 mechanical findings. Predicted mean 6.39, acceptance 79%: Likely accept. Every reviewer moved up, the skeptical senior reviewer by the most, because the claims now matched the tables.

The concerns that remain are the kind a paper carries into a real review: a human evaluation would settle the magnitude, a larger model would show scale, one non-news corpus would show domain transfer. None is blocking and each is now a known cost rather than a surprise. Open the second report.

The author submitted the revised draft. The venue decision belongs to a real panel; what PanelSim changed is that the author walked in knowing what that panel would say, with the answers already in the paper.

Accept (poster)Area chair, confidence 4 of 5

The revised submission establishes its central claim: the contrastive term, not the extra training signal, carries the faithfulness gain, and the ablation and the seed variance now show it (R1, R2). The independent judge answers the metric circularity that R1 and R5 raised on the first version, and the limitations section states the retrieval condition the negatives depend on. R7 still wants a human evaluation, which I read as a request for the camera-ready rather than a blocking issue. The reviewers agree the comparison against decoding-time interventions is the right one and that the paper adds nothing at inference. I recommend acceptance.

Reviewer by reviewer

ReviewerBeforeAfterChange
Senior researcher in faithfulness of abstractive summarization
Leans on the evidence
4.76.7+2.0
Researcher on summarization benchmarks and evaluation
Leans on the evidence
4.75.7+1.0
Researcher in contrastive learning for sequence models
Leans on the formal content
4.75.7+1.0
Industry researcher building retrieval-augmented pipelines
Leans on cost and scale
5.76.7+1.0
Senior researcher in faithful generation
Leans on novelty and positioning
4.74.70.0
Postdoc in dense retrieval
Leans on novelty and positioning
4.76.7+2.0
Senior NLP researcher outside summarization
Leans on novelty and positioning
3.75.7+2.0
PhD student working on hallucination in retrieval-augmented models
Leans on reproducibility and impact
4.76.7+2.0
Applied NLP practitioner deploying summarization
Leans on reproducibility and impact
5.76.7+1.0
Researcher in calibration of language models
Leans on the presentation
5.77.7+2.0

What is left

MinorNo human faithfulness judgments (2/10)
MinorAblation covers the negative source but not the margin form (1/10)
MinorBaseline threshold tuning still not described (1/10)
MinorMargin saturation is not analysed (1/10)
MinorLargest model tested is BART-large (1/10)
MinorComparison with a contrastive faithfulness baseline is indirect (1/10)
MinorArtifact release timing (1/10)
MinorDetection of faithful-but-wrong summaries (1/10)
MinorQualitative examples would help (1/10)
MinorStill compact for the format (1/10)

Run the same two reports on your paper.

The first report is $9.95 per paper; the revision run is $19.00 and you choose what it executes.