We ran a live Critical Reasoning workshop on assumption-family questions, hosted by
hr1212 Harsh Rumalwala. Three questions, roughly ninety minutes, and the whole session was built around one idea: think of your own answer before you read the choices, then use negation to confirm it.
This post is the full debrief. All three stimuli, the pre-thinking the room produced, what an AI produced for the same questions, and the reasoning that settled each one. Answers are behind spoilers, so you can work through them first.
THE METHODTwo steps, in this order.
1. Pre-think before looking at the options. Read the argument, find the gap, and write down what would have to be true for the conclusion to hold. Doing this before you see the answer choices is what stops you from being talked into a wrong answer by well-written distractors.
2. Negate to confirm. As Harsh put it during the session:
Quote:
For assumption, you guys might know the trick that you negate the option choice, and whichever option choice kind of breaks the conclusion, that's kind of super easy to pick it up.
Please note that the negation test is a check on a candidate answer, not a way to find one. If you negate all five choices without pre-thinking first, you will usually find two that feel like they weaken something. The pre-thinking is what tells you which gap actually matters.
QUESTION 1 - WEAKENCommercial airlines experience more flight incidents when some pilots have untreated sleep disorders than when none do. Since pilots who have previously been treated for a sleep disorder are somewhat more likely than other pilots to experience one again, any airline seeking to reduce flight incidents should permanently prohibit any pilot who has ever received treatment for a sleep disorder from flying commercial aircraft.
Which of the following, if true, most seriously undermines the argument above?- (A) Some airlines require pilots receiving treatment to take temporary medical leave.
- (B) Many flight incidents result from weather conditions rather than pilot error.
- (C) Pilots who know that receiving treatment would permanently end their flying careers often avoid reporting symptoms and continue flying without treatment.
- (D) Long-haul pilots are more likely than short-haul pilots to experience sleep-related fatigue.
- (E) Some flight incidents are caused by mechanical failures rather than by pilot error.
What the room pre-thought- Sleep disorders can be treated, so a treated pilot might actually be safer
- Pilots would stop getting treated
- Pilots would stop disclosing, they would just hide it
- Untreated pilots are the actual risk group and the ban does not touch them
- It would cause a pilot shortage
- Is the treatment even effective
What AI pre-thought for the same question- It is easy to know which pilots received prior treatment
- This policy will reduce actual incidents
- Other high-risk pilots are not being addressed
- Treatment history is a better predictor than current medical evaluation
It is interesting to note that the room found the winning idea and the AI did not. Two attendees independently landed on the disclosure problem, which is exactly what the answer turns on.
The answer is (C).
The argument's own first sentence says the incidents come from untreated sleep disorders. The proposed ban punishes exactly the people who did the right thing and got treated. (C) shows the policy is self-defeating: make treatment career-ending and pilots stop reporting symptoms, which grows the untreated population, which is the group the stimulus already blamed for incidents. The policy makes the stated problem worse.
(B) and (E) are the same trick twice. Both say some incidents have other causes, which does not matter, because the argument never claimed sleep disorders cause all incidents. (D) compares long-haul and short-haul, a distinction the argument does not make. (A) is close to relevant but temporary leave during treatment is not the permanent ban being argued for.
QUESTION 2 - STRENGTHENA multinational company's network logs reveal a pattern of unauthorized access and encrypted files characteristic of a ransomware attack. Cybersecurity analysts hypothesize that the breach resulted from a well-documented ransomware campaign that targeted organizations worldwide in May 2024.
Which of the following, if true, most strongly supports the analysts' hypothesis?- (A) The company's servers contained software versions that were commonly used both before and after May 2024.
- (B) No system logs generated after May 2024 were found on the compromised servers, whereas logs from before that month were recovered in abundance.
- (C) Numerous cybersecurity reports confirm that a global ransomware campaign occurred in May 2024.
- (D) Several employee laptops were running operating systems released during 2023.
- (E) Backup files created in June 2024 were recovered intact from the compromised servers.
What the room pre-thought- Find evidence that the log pattern matches what was seen in the May 2024 campaign
- Rule out a different, similar ransomware campaign
- Pull observations from another company hit by the same attack
- Show this campaign actually reached this company
What AI pre-thought- There was not another campaign that could have caused it
- The company was actually exposed to that campaign
- The characteristics match that particular campaign, not ransomware in general
Both lists converge on the same three jobs: match the signature, place the company in the blast radius, and eliminate alternatives. That convergence is a good sign your pre-thinking is sound.
The answer is (B).
Riya Mondal made the argument that settled it: logs stop being generated after May 2024 while earlier logs survive in abundance. That timestamp boundary places the disruption in the same month as the campaign, which is the link the hypothesis needs.
(C) is the trap, and several people in the chat went for it. It confirms the campaign existed, which nobody disputed. The stimulus already treats the campaign as well-documented. Confirming a premise adds nothing. The gap is not whether the campaign happened, it is whether this breach came from that campaign.
(E) actively hurts the hypothesis. Intact backups from June suggest the systems were still functioning after May. (A) and (D) are date-flavoured noise with no causal link.
QUESTION 3 - ASSUMPTION (the one that split the room)Human activity has been releasing chlorofluorocarbons (CFCs) into the atmosphere for many decades, and these compounds contribute to the depletion of the Earth's ozone layer. Therefore, by measuring the average rate at which the ozone layer has thinned over the past hundred years, and then calculating how many centuries of such thinning would have been required for the ozone layer to decline from a hypothetical original state to its current thickness, scientists can accurately estimate the maximum age of widespread ozone depletion.
Which of the following is an assumption on which the argument depends?- (A) The amount of CFCs released into the atmosphere during the past hundred years has not been unusually high compared with earlier centuries.
- (B) At any given time, all regions of the Earth's atmosphere contain about the same concentration of ozone.
- (C) Certain naturally occurring gases also contribute to ozone depletion.
- (D) No method other than one based on ozone thickness provides a more accurate estimate of the maximum age of widespread ozone depletion.
- (E) None of the ozone destroyed in the atmosphere is regenerated through naturally occurring chemical reactions.
What the room pre-thought- Depletion has happened at a uniform rate in the past
- The ozone layer is not rejuvenating through some other mechanism
- No other factor is contributing to depletion
- Confidence in the average rate and the calculation itself
What AI pre-thought- Constant rate assumption
- No offsetting process
- Other causes of thinning
- Measurement accuracy
This is where it got interesting. Most of the room, and the AI, put "no offsetting process" high on the list. That maps to (E). Several attendees argued for (E) out loud.
The answer is (A). Harsh confirmed it after a long back and forth.
The argument takes the average rate from the last hundred years and projects it backwards. That only works if the last hundred years were representative. Negate (A): CFC release over the past century was unusually high compared with earlier centuries. Now the recent rate is faster than the historical rate, so dividing total depletion by that inflated rate returns too few centuries. The method underestimates the age, and the conclusion that it "accurately estimates the maximum age" collapses.
Why (E) is so tempting and still wrong. This is the takeaway worth keeping.
(E) is stated at an extreme, "none of the ozone destroyed is regenerated". Negate it and you get "some ozone is regenerated". rajiv joarder worked through the direction of that error in the session: if some ozone comes back, the true depletion was larger than measured, so the estimate would come out too high.
Now look at what the conclusion actually claims. It claims a maximum age. An estimate that runs too high is still a valid ceiling. An estimate that runs too low is not. So negating (E) bruises the argument, while negating (A) breaks it.
That asymmetry is the whole question. When a conclusion is phrased as a maximum, a minimum, "at least", or "no more than", work out which direction of error is actually fatal before you pick between two choices that both feel relevant.
Quick notes on the rest. (B) is about regional distribution, and Rishika pointed out the argument runs on averages, so regional variation does not touch it. (C) gives an additional cause of depletion, which is not something the argument needs to be true. (D) compares this method against other methods, and the argument never claims to be the best method, only an accurate one.
THREE THINGS TO CARRY OUT OF THIS- Pre-think first, every time. On question 1 the room produced the correct idea before seeing a single answer choice. That is what makes (C) obvious instead of tempting.
- Negation confirms, it does not discover. Two choices will often both wobble the argument under negation. Only one breaks it. Knowing your pre-thought gap tells you which.
- Read the conclusion's direction. "Maximum age" was doing more work in question 3 than the ozone chemistry was. Extreme-sounding options like "none" and "all" are not automatically wrong, but check whether the error they introduce actually points the fatal way.
Thanks to everyone who spoke up,
Rishika,
rajiv joarder,
Nitesh Motwani,
Riya Mondal and
Varundeep Singh, and to
Harsh Rumalwala for running it.
Post your reasoning below, particularly if you picked (E) on the last one. Working out why the near-miss is a near-miss is worth more than getting it right by instinct.