What Should the Human Part of Science Become?

1. The change is already inside the field

I have been thinking about AI and research less as a distant forecast and more as a change in the texture of an ordinary working day. A calculation that once took an afternoon can be checked in minutes. A counterexample that might have emerged after several conversations can appear in the first serious exchange with a model. Code can be written, tested, and changed while the mathematical idea is still forming. Literature from a neighboring area can be entered without first learning all of its local shorthand. None of these uses is completely reliable, but together they change what it feels like to start and pursue a research problem.

Two recent talks helped sharpen the question for me. Hsin-Yuan (Robert) Huang framed the automation of research as an open problem for researchers themselves: what should remain valuable, and how should scientific practice respond when parts of it can be automated [1]? Terence Tao’s ICM lecture asked a related question for mathematics, while keeping the emphasis on the human purposes of the subject [2]. Tanya Klowden and Tao place the same concern in a wider philosophical setting and argue for a human-centered use of AI [3]. I found the tone of these discussions more useful than either celebration or panic. The technology is moving too quickly for confident predictions, but the questions are no longer hypothetical.

For quantum information, QIP makes them concrete. QIP 2027 will take place in Singapore from February 20 to 26, 2027; talk registration closes on September 28, 2026, and completed talk submissions are due on October 5 [4]. The conference has already published rules for generative AI. Use of these tools is not, by itself, counted either for or against a submission. Authors must confidentially indicate substantive use to the relevant chairs, cannot list a model as an author, and remain responsible for correctness, originality, attribution, and integrity [5]. These are sensible first steps. They also show how quickly a philosophical discussion has become a practical policy problem.

QIP is not merely a place where correct theorems are announced. It is a week in which a community decides what deserves shared attention. Its program influences what students learn, what problems appear central, and which kinds of contribution become legible as excellent work. If AI increases the supply of technically polished results, the program committee will not only be filtering correctness. It will be deciding what is worth asking hundreds of people to understand together.

Several arguments in this essay grew out of short public notes that I posted between August 7 and 12, 2026, and from the discussions that followed. The recurring question was simple: if correct results become abundant while the time to read and digest them does not, what becomes scarce [6]? The replies made the problem less abstract. They raised experiments, explanation, student training, language, access, credit, and the possibility that our current unit of contribution, one PDF attached to a list of authors, may not survive unchanged.

2. What AI has actually done in quantum information

It is useful to begin with examples, because otherwise the conversation slips too easily between routine language editing, an automated proof search, and an imagined autonomous scientist. These are not the same activity. They create different questions about reliability and credit.

2.1 A proof experience close to home

In a recent preprint with Rishikesh Gajjala and Tobias Haug, we study the two-copy distillability problem for Werner states [7]. A bipartite entangled state is distillable if local operations and classical communication can extract useful near-maximally entangled pairs from copies of the state. Werner states form a highly symmetric family, but the finite-copy distillability question has resisted a complete understanding for a long time. Our preprint claims a sharp two-copy characterization through a dimension-free partial-trace inequality.

The initial proofs were generated by ChatGPT 5.6 Sol. We then verified and rewrote the arguments, reorganized the presentation, and added context. I shared the experience publicly because hiding the workflow would have made the scientific record less informative [8]. The result is currently a preprint, not a peer-reviewed fact, and that distinction matters. Yet the episode already changed my view of what counts as a realistic AI contribution. The model was not merely polishing English or expanding algebra that we had already supplied. It participated in the proof process in a substantive way.

The public discussion around that post was revealing. Felix Huber reported that he, Thomas Fraser, Balazs Pozsgay, and Istvan Vona had independently obtained a GPT-assisted proof and became confident only after careful checking. Chiranjib Mukhopadhyay argued that useful results should not be ignored because AI contributed, while insisting on attribution and pointing to the review burden that a flood of such work may create [8]. Both comments place responsibility in the right location. A generated argument is not yet a trusted result. The work of checking, reconstructing, comparing with the literature, and making the idea intelligible remains real scientific labor.

2.2 Unclonable encryption and unstable priority

A second example comes from quantum cryptography. Unclonable encryption uses the no-cloning structure of quantum theory to encode a classical message so that an adversary cannot split a quantum ciphertext into two systems that both remain useful after the key is revealed. In July 2026, Prabhanjan Ananth and Amit Sahai posted an efficient, information-theoretically secure one-bit construction in the plain model [9]. Their paper states that the construction and main proof ideas were generated by Codex using GPT 5.6, while the human authors refined and verified the claims and accept responsibility for them.

Seyoon Ragavan posted a concurrent and independent construction based on random Pauli eigenstates [10]. That paper gives an explicit account of the interaction: the model explored several approaches, the author pushed it toward a simpler spectral proof, and the final work included a Lean formalization of the main correctness and security claims. The two works found essentially the same central object at nearly the same time.

This is scientifically exciting, but it also exposes the weakness of a credit system built mainly around temporal priority. If several strong systems can search a space of ideas in parallel, being first may depend on access, budget, orchestration, and a narrow difference in posting time. The durable contribution may lie elsewhere: who identified the right problem, who understood the proof, who found the clean abstraction, who checked the security definition, who connected the result to earlier work, and who made the method reusable. Priority does not disappear, but it becomes a noisier proxy for scientific contribution.

2.3 Discovery is not confined to proofs

Quantum optics has provided an earlier and more mature form of computer-assisted discovery. The MELVIN system searched combinations of optical components and proposed experiments for producing complex entangled states [11]. PyTheus later turned this direction into an open-source graph-based framework and reported one hundred designs covering states, measurements, communication protocols, and gates [12]. These systems do not make an experiment happen by returning a graph. A physicist must still recognize whether the design is meaningful, translate the abstract components into laboratory constraints, control loss and distinguishability, and decide whether the resulting phenomenon teaches us anything. Still, the search system can expose a configuration that a human would not have tried.

This matters for the usual theory-versus-experiment picture of automation. Experimental science is protected from prompt-level automation by matter itself. Photons are lost, devices drift, fabrication fails, and calibration consumes time. But experiment design, control, data analysis, error diagnosis, and the search over candidate configurations can all be assisted. The human role does not remain safe merely because a laboratory contains hardware. It changes in a different way.

2.4 Formal infrastructure is arriving

The other important development is formal verification. Lean-Quantum has begun building a basis-independent Lean 4 library for finite-dimensional quantum information. As a demonstration, it formalizes the dataprocessing inequality for the sandwiched Renyi relative entropy and provides reusable interfaces for states, channels, partial traces, Choi operators, and Stinespring representations [13]. A separate 2026 preprint introduced two benchmarks containing 76 Lean theorem-completion tasks from quantum algorithms and quantum information [14]. The reported model scores leave substantial room for improvement, but the important point is that quantum-information proofs are becoming a measurable target for agents operating inside a proof assistant.

These examples sit at different levels of maturity. The optical-design papers include peer-reviewed work. Lean-Quantum and the theorem-proving benchmarks are recent preprints. The Werner-state and unclonableencryption claims are also recent preprints and must be scrutinized accordingly. A public disclosure of AI use does not make a theorem correct, and a dramatic workflow does not substitute for peer review. The responsible conclusion is narrower: AI-assisted discovery in quantum information is real enough that our norms can no longer be designed for a world in which models only improve grammar.

3. When getting an answer becomes cheap

There is a tempting way to divide the future: machines will generate and verify, while humans will choose questions and interpret. It may be a useful description of the present, but it is not a stable definition of the human role. Models may become better at proposing questions, finding abstractions, and writing explanations too. If we define human value as whatever today’s system cannot do, we will be forced to retreat each time the capability boundary moves. The human role should instead be stated in terms of responsibility. People and institutions must decide which questions merit resources, which formalization preserves what matters, when evidence is strong enough to call a claim knowledge, and what consequences follow from acting on it. These duties remain ours even if a machine becomes very good at helping with them.

Tao has separated mathematical work into proof generation, proof verification, and proof digestion. The distinction emerged vividly in a recent project on primitive sets: a method suggested by model output led to new results, but substantial human work was required to understand the method, place it in the literature, and develop its consequences [15, 16]. I find this separation useful for quantum information as well.

Proof generation asks for a path from assumptions to conclusion. Verification checks that the path is valid. Digestion asks why the path works, which parts are reusable, what the theorem actually says, and what should be done next. Our current publication culture often compresses all three into a single output and then uses that output as a proxy for contribution. AI begins to pull them apart.

Consider a security proof. A model may find the operator inequality that closes the argument. A proof assistant may certify each logical step. Yet a researcher still has to explain why that inequality captures an adversary’s optimal strategy, whether the security game is the right one, which restriction makes the result possible, and whether the construction survives composition with another protocol. In error correction, an automated search may return a code with excellent finite-size parameters. Someone must still understand the noise model, decoding assumptions, syndrome-extraction circuit, and hardware geometry before calling it useful. In quantum algorithms, a formal speedup may disappear when state preparation, precision, or readout is counted. The technical answer is important, but it is not the whole scientific answer.

If proof production becomes cheap, difficult proofs do not become worthless. Difficulty was never the final reason a theorem mattered. We care because a theorem changes the boundary of what is possible, clarifies a phenomenon, supplies a method, or reorganizes how we see a subject. A short argument can be more important than a long one. The arrival of AI makes this old fact harder to ignore.

There is a risk here. We may respond to cheaper proof production by demanding more theorems per paper, more experiments per theorem, or a faster publication cycle. That would preserve the appearance of scarcity by raising the quantity threshold. It would also move us toward a literature that almost nobody can read. A healthier response is to value the work of selecting, compressing, and explaining. The scarce achievement may be a five-page argument that tells the community what one should learn from fifty generated lemmas.

4. A correct proof is not the same as the right theorem

Formal verification will be essential in a world of long machine-generated arguments. Lean checks proof terms with a small trusted kernel, making the logical dependencies explicit [17]. This is a much stronger certificate than an assurance that an appendix was checked carefully. It can expose missing hypotheses, type mismatches, and appeals to lemmas that do not say what the prose claims. Shared formal libraries also make results reusable rather than leaving each paper to rebuild its own foundations.

But formalization moves part of the trust problem rather than eliminating it. A proof assistant proves the statement we give it. Someone must still check that the formal statement represents the scientific claim we care about. I made this point in a public note about Lean: correctness and understanding are distinct, and science needs both [18].

Quantum cryptography makes the distinction sharp. Much of the content of a result is carried by the definition: Is the adversary computationally bounded or information-theoretic? Does it receive quantum side information? Can it choose messages adaptively? Are two devices isolated, and at what stage? Is security standalone or composable? A perfectly formal proof of a weak or misaligned definition may be mathematically correct and operationally irrelevant. The proof assistant cannot infer the intended threat model from the surrounding motivation.

The same issue appears across quantum information. A resource theory depends on which operations are free. A statement about an advantage depends on the classical baseline and on which costs are counted. A metrological claim depends on the estimator, prior information, repetitions, and resource accounting. An error-correction threshold depends on a noise model and a decoder. A theorem about a channel may be correct while the physical process of interest is not well approximated by that channel. In each case, formalization certifies an implication inside a model. It does not certify that the model is the right interface with nature.

This is not an argument against formal methods. It is a reason to build them carefully. A trusted library should contain more than technically convenient definitions. Its interfaces should be discussed by the community, connected to operational tasks, and accompanied by examples and counterexamples that show what they include and exclude. Formalization can improve scientific understanding when the act of making assumptions explicit forces us to resolve ambiguity. It becomes dangerous only when the machine-checked badge discourages further thought about what was checked.

There is also a difference between verifying and explaining. A proof term may be correct but unreadable. An automatically generated derivation may use a large library in a way that gives no human clue about the central mechanism. In such cases, we possess a certificate but not yet a scientific account. The next valuable contribution may be a new lemma, picture, or operational interpretation that makes the proof obvious in retrospect. That work should receive credit.

5. What should scientific credit measure?

Divesh Aggarwal put the difficulty plainly in response to my note on future QIP formats: if something previously regarded as a QIP-level paper can be generated in a day, does credit still mean anything [19]? I think it does, but paper count will become an even weaker measure of it.

Credit has never tracked effort perfectly. A theorem can follow from one lucky observation or ten years of failed attempts. Experimental results may depend on infrastructure built by people whose names are distant from the headline. Senior researchers may receive more recognition than junior collaborators for the same idea. Gift authorship already stretches the connection between the author list and the work. AI does not create these problems, but it removes a convenient fiction: that the technical content of a paper approximately records the intellectual labor of its authors.

We should separate at least three questions. First, is the result correct and important? Second, what did each human contributor do? Third, how did an AI system enter the process? The first question belongs to scientific evaluation. The second and third belong to attribution and credit. Mixing them too early creates bias; ignoring the latter two makes the record misleading.

When a model supplies the central proof idea, human contribution may still be substantial. A researcher chose or recognized the problem, built the context, designed the harness, rejected unpromising routes, checked the output, repaired gaps, connected the result to the literature, wrote an intelligible proof, and accepted responsibility. But these activities should be described, not silently converted into the claim that the model was only a tool like a calculator. The comparison obscures more than it clarifies. A calculator evaluates an expression chosen by the user. A research agent may propose the expression, the lemma, or the whole route.

Authorship should remain human because authorship carries duties that a model cannot discharge. An author must answer questions, correct errors, disclose conflicts, respond to criticism, and bear reputational responsibility. This is consistent with QIP 2027’s rule that generative systems cannot be authors [5]. It does not follow that the model’s role should disappear. A concrete contribution statement can record whether AI was used for writing, search, code, conjecture generation, counterexamples, proof strategy, formalization, or experimental design.

Percentages such as “the idea was 60 percent human” will not help. They are impossible to calibrate and invite strategic reporting. A better record is factual: the model and access date, the broad workflow, the stage at which a key object appeared, the checks performed, and the artifacts that can be shared. Omar Shehab has argued publicly for provenance or a record of effort alongside correctness in AI-assisted science. That proposal is attractive, provided the record does not become a demand to publish every private conversation or every failed prompt. Research requires space for exploratory thought. Provenance should make a contribution legible, not turn intellectual work into surveillance.

The hardest problem is junior credit. Students and postdocs need identifiable contributions in order to get jobs and build independence. Replacing paper-centered talks entirely with broad surveys would favor researchers who already possess visibility. At the same time, pretending that every AI-assisted theorem reflects the same kind of individual accomplishment will eventually make the credential meaningless. Near-term credit may need multiple signals: a clear contribution statement, evidence of command of the work, talks or defenses that reveal understanding, software and formal artifacts, and a record of questions pursued over time. None is perfect. Together they are better than counting PDFs.

One proposed response is a hard cap on submissions per person. This could reduce flooding and discourage attaching senior names to many papers, but it is blunt. It can also punish genuine collaboration and shift strategic behavior into the author list. Another response is to give more space to experimental work, long-running projects with a visible history, or problems that remain difficult for the combined class of humans using current language models, which one might denote informally by HumanLLM. These may help in the near term. They are not stable definitions of value. Experimental design is also being automated, project histories can be gamed, and the boundary of HumanLLM will move.

The deeper answer is that credit should track understanding, judgment, responsibility, and contribution to a shared scientific structure, not merely possession of an answer. We will not obtain this from one metric.

6. Let the reviewer think before the model does

AI can help peer review. It can check algebra, search for a missing case, compare a claim with public literature, generate small test instances, or help a reviewer navigate a long appendix. Reviewers are already overloaded, and the number of submissions may rise sharply. Refusing all assistance is unlikely to be sustainable.

I am much less comfortable with a model supplying the scientific judgment. When authors receive a report, they want more than a binary check of the proof. They want to know how another researcher sees the work. Is the question natural? Does the abstraction organize other problems? Is the regime meaningful? Did the reviewer notice a connection the authors missed? That perspective is part of peer review, and it cannot be outsourced without changing what review is.

A useful default is human first, AI second. The reviewer should read the abstract and introduction, inspect the main result, and form an independent view before asking a model for a summary or critique. I argued for this ordering in a public note on AI-assisted review [20]. The order matters because a summary is also a frame. It directs attention toward some contributions and away from others. Once a model has described a paper as incremental, elegant, or confused, it becomes harder to encounter the work freshly.

After that first reading, tools may be used for bounded checks. The human reviewer should verify any alleged error directly, consult the cited source rather than trust a generated comparison, and write the report in their own voice. A reviewer should never paste confidential discussion, other reviews, or identifying information into an external service without authorization. At QIP 2027 the technical manuscript must be public on arXiv, which reduces one confidentiality concern, but program-committee deliberations and reviewer text remain confidential [5]. Public availability of a paper is not permission to export the rest of the review process.

Disclosure poses a related problem. It is useful to know where AI entered a research workflow, especially while norms are changing. Yet if a reviewer sees “heavily AI-assisted” before reading the work, the label may change how surprising or worthy the same theorem feels. I previously suggested separating disclosure from evaluation for this reason [21]. QIP 2027 has adopted a closely related structure: substantive use is disclosed to the program and topic chairs, not made public through the submission field, and is formally neutral to acceptance [5]. This arrangement should generate information without turning disclosure into a warning label.

No policy will work if honesty is punished. If authors believe disclosure lowers acceptance, hiring, or grant prospects, the most careful researchers will be more transparent and may be penalized for it, while others remain silent. A successful policy must make truthful disclosure ordinary, specific, and low-stigma. It should also enforce responsibility: an author cannot answer a criticism by saying that the model wrote that part.

7. Why being stuck still matters

Research involves a remarkable amount of being stuck. We try an idea, find a counterexample, weaken an assumption, and discover that the weaker statement is still false. We talk to someone and realize that we misunderstood the problem. From the outside, this looks inefficient. From the inside, it is often how intuition is built.

If a model immediately gives the right counterexample or proof strategy, we save time. That is valuable. But something can also be lost, especially in training. The answer is not to preserve every old difficulty. There is little educational value in forcing a student through a long routine calculation that software can check in seconds. The useful distinction is between unnecessary friction and productive struggle [22].

Unnecessary friction consumes time without changing how the student thinks. It includes searching for a missing minus sign, translating a standard identity between conventions, or writing boilerplate code. Productive struggle reveals structure. It teaches which assumptions matter, which examples are dangerous, what scale an answer should have, and why a tempting route fails. The two can look similar while they are happening, which is why a universal rule such as “use AI immediately” or “do not use AI” is too simple.

Lennart Binkowski asked in the public discussion how to classify the textbook derivation of the hydrogen-atom wavefunction: is it insight or tedium? Rajarsi Pal emphasized the intrinsic reward of finding one’s own route and suggested that intuitive explanation may become a particularly human contribution [22]. These comments capture the educational problem. The answer depends on what the student is meant to learn. Solving a differential equation by hand may be routine in one course and conceptually central in another.

In quantum information, a student may ask a model to prove monotonicity of a divergence, derive a decoder update, or find a cheating strategy for a protocol. The adviser should ask what competence is being developed. If the goal is to learn operator inequalities, reading the generated proof too early can remove the very experience that builds judgment. If the goal is to test a new cryptographic definition, automating standard matrix manipulations may free the student to focus on the adversary model.

A reasonable practice is staged use. Let the student first state what they believe, try a few routes, and record why they fail. Then use the model as a critic, adversary, or source of alternatives. Finally, require the student to reconstruct the successful argument without leaning on the conversation. The point is not to slow research down for moral reasons. It is to make sure that speed does not replace learning.

Training must also change at the level of evaluation. A take-home proof that can be generated instantly no longer reveals much. Oral explanation, live problem formulation, error diagnosis, and the ability to modify an argument under a changed assumption may become more informative. Students should learn formal verification and AI-assisted workflows, but also how to work when the tool fails. Intellectual independence now includes knowing when not to trust an apparently excellent answer.

8. AI can narrow one gap and widen another

AI can lower barriers in a real way. A researcher entering a new field can ask basic questions without embarrassment, get help with unfamiliar notation, explore literature, write small pieces of code, and fill gaps in background knowledge [23]. This can make interdisciplinary work easier. It may help talented researchers outside established groups enter areas that once depended heavily on local expertise and personal connections.

Victor V. Albert pointed to another inclusion benefit in response to my later post: researchers who are not fluent in English, and nonexperts with a correct result, may be able to communicate it more clearly so that the community can inspect the substance [24]. Scientific English has long acted as a gate. Reducing that gate is not cosmetic. A good idea that cannot be presented in the expected style is often not evaluated fairly.

At the same time, a new inequality may appear. The advantage will not belong only to people who can afford a subscription. It may belong to groups with sustained access to the strongest models, large API budgets, private tools, compute, and enough time to build effective research harnesses. There is a large difference between occasionally asking a chatbot a question and operating a system that launches many agents, supplies curated libraries, tests outputs automatically, and learns from repeated failures. Even if model prices fall, workflow fluency and private context may remain unequal.

Reputation can compound the difference. If routine problem solving becomes cheaper and value moves toward choosing directions and attracting attention, an established researcher may find it easier to make the community care about a generated result. A junior researcher outside the main network may have the result but not the audience. Automation of production does not imply democratization of recognition.

This is why open models, shared formal libraries, university access programs, transparent workflows, and public benchmarks matter. Shared infrastructure will not erase inequality, but it can keep the basic ability to participate from becoming a private asset. The community should also resist evaluating researchers by raw AI-amplified output. Otherwise the groups with the largest budgets will appear the most creative simply because they can search the largest tree.

9. Negative results become more valuable, not less

The usual publication system rewards a positive theorem, a faster algorithm, or a successful experiment. Failed approaches and null outcomes are harder to publish, even when they prevent others from repeating expensive work. Publication bias is known to distort the literature and waste resources by sending researchers back into unproductive directions [25]. The AI era makes this problem more severe.

AI can generate plausible directions much faster than humans can check them. It can propose hundreds of ansatze, reductions, code constructions, or experimental configurations. Verification remains costly. If only successful paths are reported, many groups may spend their limited attention falsifying the same seductive idea. A careful negative result becomes a map of where not to search. It saves compute, laboratory time, and, most importantly, human thought.

Quantum information already has a strong culture of no-go theorems. No-cloning is a negative statement that created quantum cryptography. Eastin-Knill-type limitations shape fault-tolerant architecture. Lower bounds, impossibility results, and separations tell us which resource is doing the work. A negative result is not a failed paper when it identifies a genuine obstruction under natural assumptions.

One example from my own work is the study with Tobias Haug and Dax Enshan Koh on pseudorandom quantum states and unitaries [26]. Pseudorandom objects are efficiently generated but appear Haar-random to efficient observers, making them useful across complexity, cryptography, and many-body physics. We proved, among other limitations, that the relevant pseudorandomness is not robust to non-negligible noise, ruling out the desired form on noisy intermediate-scale and early fault-tolerant devices. I described the work publicly at the time as a collection of negative results [27]. Its value lies precisely in closing tempting routes and clarifying the resources that pseudorandomness requires.

Not every failed prompt deserves a paper. A useful negative result needs a clearly defined claim, natural assumptions, evidence or proof strong enough to exclude a meaningful region, and an explanation of what remains open. A log saying that one model did not find a solution is usually evidence about the workflow, not about the mathematics. The distinction is between failure to discover and discovery of a limitation.

Conferences and journals could make these contributions easier to see. A short, well-checked no-go result may save more community resources than another modest positive construction. Reproducible counterexamples, benchmark failures, and carefully delimited impossibility statements should be judged by the territory they remove from the search space. When generated possibilities are abundant, knowledge of impossibility becomes unusually scarce.

10. The real bottleneck is attention

If AI greatly increases the number of correct results, human attention will not increase at the same rate. The supply of theorems may grow much faster than the time available to read, understand, and discuss them. A result does not become worthless because AI made it easy to find. If the answer is important, it remains important. The institutional question is different: what is worth spending human time understanding [6]?

This question is uncomfortable because attention is not allocated neutrally. Fashion, reputation, institutional prestige, and presentation already matter. AI may make the allocation problem worse by producing more polished claims than reviewers can inspect. It may also help by summarizing, connecting, and filtering. The danger is a closed loop in which models trained on what was previously prominent recommend more of the same, while unusual questions disappear because they resemble nothing already rewarded.

Anurag Saha Roy suggested in the discussion that theoretical work with interesting or implementable experimental consequences may retain attention as the number of theoretical results grows [6]. That seems plausible. Physical realization imposes a cost and supplies a form of contact with nature that a generated proof does not. But it cannot be the only filter. Some of the deepest work in quantum information is valuable long before an experiment is possible, and some beautiful theory tells us that an experiment cannot do what we hoped.

The better response is plural. We need people who prove, people who verify, people who build, and people who digest. We need surveys that organize a direction, short notes that isolate a decisive obstruction, formal libraries that make results reusable, software that allows others to test an idea, and talks that explain why a theorem matters. The paper will remain important as an archival object, but it need not remain the only unit of contribution.

There is a further reason to protect human attention. Attention is not just a channel for distributing credit. It is where communal understanding is formed. A field advances when people share abstractions, argue about definitions, notice that two problems are the same, and decide that an apparently technical caveat is actually central. If every researcher privately queries a model and receives a different stack of results, the field may produce more while knowing less together.

Conferences therefore become more important, not less, but their purpose shifts. They are occasions for compression and common orientation. Selection should ask what an audience will understand differently after the talk. That is a higher standard than technical correctness, and it is unavoidably a matter of judgment.

11. What a QIP-like community can try now

The QIP charter says that the conference should represent the preceding year’s best research and ordinarily selects only a small number of contributed talks [28]. This scarcity is part of its value. The aim should not be to preserve an old ritual at any cost, but to keep the program scientifically useful while the production process changes.

I would begin with modest rules whose purpose is clear.

Keep evaluation neutral to tool use. An important AI-assisted result belongs in the scientific conversation. A permanent human-only track would become difficult to define and would exclude work that the community may need to understand. The goal is a strong program, not a moral test of which tool was used. QIP 2027’s explicit neutrality is therefore the right baseline [5].

Ask for concrete, staged disclosure: Authors should identify substantive use by category rather than assign percentages. The disclosure should first be visible only to people responsible for policy and conflicts, after ordinary reviewers have evaluated the content. This protects transparency while reducing framing bias. For unusually substantive cases, a public statement in the arXiv paper is appropriate. A minimal reproducibility record might include the model, date, broad harness, and human verification steps. Exact prompt transcripts should be optional unless they are necessary to reproduce a claim.

Make understanding part of presentation. The human speaker should be able to explain the main idea, defend the assumptions, answer questions, and say what remains uncertain. Barbara Terhal has publicly suggested short videos as evidence of how the humans responsible for a result communicate it, regardless of how it was generated [29]. A video requirement may create its own language and accessibility biases, so it should not become a polished performance contest. The underlying principle is sound: a live scientific meeting should select not only a theorem but an act of human explanation.

Protect the route by which junior researchers receive credit. I have argued that QIP may eventually include more talks about a direction rather than one technical result, but that a mixed format is safer than replacing paper-centered talks [19]. Invited synthesis talks can help digest an abundant literature. Contributed talks should continue to give students and postdocs a visible platform for work they understand and led. Contribution statements, presenter choice, and questions during the talk can help. The current QIP student-paper rule, which asks whether students completed a significant majority of the work and key ideas, will need careful interpretation when a model supplies part of those ideas [5]. Simply ignoring the model will not resolve the question.

Permit AI assistance in review only within boundaries. Reviewers should form an independent view first, use tools for bounded checks, verify alleged errors, and own the final report. They should not export confidential material. Chairs should state these rules explicitly so that reviewer practice does not become an invisible variable. If models are used to check public manuscripts, the report should still tell the authors what a human expert concluded.

Give negative and expository work a real path. If attention is the bottleneck, a rigorous negative result or an exceptional synthesis can be more valuable than an additional isolated theorem. This does not require lowering standards. It requires judging usefulness in the right unit. A negative result should close a meaningful region; an exposition should produce understanding that the component papers do not provide. Both should identify the people who did the intellectual work.

Collect evidence and revise. The first policies will be wrong in places. Conferences should collect aggregate information about how tools were used, how often disclosures were substantive, whether reviewer load changed, and where disputes arose. Policies should have an explicit review date. A rule written for QIP 2027 should not quietly become permanent after the systems and workflows have changed.

These proposals do not answer every hard case. Suppose a model produces a theorem, a proof, a Lean certificate, and a clear draft with minimal human input. Should the person who typed the prompt receive a talk? My instinct is that the result may deserve a place while the individual credit claim remains thin. A conference selects science and speakers at the same time, so it cannot avoid the tension. It may choose a speaker who can genuinely digest and teach the result, while the provenance record prevents that act of explanation from being mistaken for sole discovery.

There will also be pressure to infer human effort from private records, timestamps, or message histories. We should resist building a bureaucracy in which researchers must prove that they suffered enough. Effort is not scientific value. Provenance matters because it supports attribution and trust, not because a long prompt log is morally superior to a short one.

12. A future worth aiming for

There is an easy pessimistic story in which machines prove theorems, output explodes, well-funded groups gain the best tools, and humans slowly lose their intellectual role. I wrote a version of this contrast in a public note, and the response reminded me that the optimistic path has several dimensions [24]. AI may remove shallow bottlenecks and allow us to spend more time on deeper ones. Researchers may ask more ambitious questions because exploring an idea becomes cheap. Formal verification may make correctness less fragile. People may move between fields more easily. Non-native English speakers may be judged more on their science and less on fluency.

Perhaps papers will become less about displaying how much technical labor went into a result and more about helping others understand what was learned. Conferences may care more about insight than proof length. We may give more credit to clear explanations, useful software, formal libraries, negative results, and well-chosen questions. None of this follows automatically from better models.

The opposite equilibrium is also plausible: more papers, faster reviews, weaker training, concentrated access, and less shared understanding. Institutions may use AI mainly to increase throughput. Researchers may feel compelled to generate continuously because everyone else can. Reviewers may delegate judgment because there is too much to read. Students may assemble impressive outputs without acquiring the ability to tell when an assumption is wrong. A field can become more productive by its metrics while becoming intellectually thinner.

We still have some choice about the norms we build. The right goal is not to reserve science for humans by keeping machines weak. It is to use strong tools in ways that enlarge human understanding and participation. This requires incentives, not only advice. If hiring counts papers, researchers will produce papers. If conferences reward digestion, careful negative results, transparent workflows, and excellent explanation, more people will invest in them.

13. Where this leaves us

AI is already doing parts of quantum-information research. It can generate proof ideas, search experimental designs, write formal code, and accelerate the movement between conjecture and test. The question is no longer whether machines will participate in science. The more important question is what the human part of science should become as that participation grows.

My answer is provisional. Humans should remain responsible for choosing questions worth asking, defining the operational problem, checking that formal statements express the intended science, interpreting the result, explaining it to others, and deciding what deserves collective attention. We should preserve productive struggle in training while automating pointless friction. We should welcome broader access while building shared infrastructure that prevents science from becoming pay-to-play. We should record substantive AI contributions without turning disclosure into stigma. We should value negative results because they conserve the resource that will become hardest to replace: careful human attention.

For QIP 2027, the first policy steps are already visible. They will not be the final answer, and they should not be expected to carry the entire burden. What matters is the direction: useful tools without surrendered judgment, formal correctness without semantic complacency, faster research without weaker researchers, and more scientific output without less science. 

Acknowledgments. I am grateful to Divesh Aggarwal, Victor V. Albert, Lennart Binkowski, Felix Huber, Chiranjib Mukhopadhyay, Rajarsi Pal, Anurag Saha Roy, Omar Shehab, and others whose public comments and questions sharpened parts of this discussion. 

 Drafting and AI-use note.  AI was used to polish a draft written in LaTeX. AI also helped generate the HTML code embedded here by the author.

References

[1] H.-Y. Huang, “How to Respond to the Automation of Research,” Simons Institute for the Theory of Computing, July 24, 2026. https://simons.berkeley.edu/talks/hsin-yuan-huang-caltech-2026-07-24.

[2] T. Tao, “Mathematics in the Age of AI,” public lecture, International Congress of Mathematicians, July 24, 2026. https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.pdf.

[3] T. Klowden and T. Tao, “Mathematical methods and human thought in the age of AI,” arXiv:2603.26524 (2026). https://arxiv.org/abs/2603.26524.

[4] Quantum Information Processing Conference, “QIP 2027,” accessed August 13, 2026. https://qipconference.org/2027/.

[5] Quantum Information Processing Conference, “QIP 2027 Call for Submissions,” accessed August 13, 2026. https://qipconference.org/2027/call/.

[6] K. Bharti, “If AI greatly increases the number of correct results,” LinkedIn post and public discussion, August 7, 2026. https://www.linkedin.com/posts/kishor-bharti-73824357_if-ai-greatly-increases-the-number-of-correct-activity-7491462490237755394-Jgq1.

[7] K. Bharti, R. Gajjala, and T. Haug, “Two-copy nondistillability of Werner states: sharp partial-trace inequalities and finite-copy extensions,” arXiv:2607.24479 (2026). https://arxiv.org/abs/2607.24479.

[8] K. Bharti, “Here’s an interesting experience we (with Rishikesh Gajjala and Tobias Haug) recently had with ChatGPT 5.6,” LinkedIn post and public discussion, July 25, 2026. https://www.linkedin.com/posts/kishor-bharti-73824357_heres-an-interesting-experience-we-with-activity-7486773460510179329-7ZJ7.

[9] P. Ananth and A. Sahai, “Unconditional Unclonable Encryption,” arXiv:2607.21551 (2026). https://arxiv.org/abs/2607.21551.

[10] S. Ragavan, “Efficient Unclonable Encryption from Pauli Eigenstates,” arXiv:2607.21811 (2026). https://arxiv.org/abs/2607.21811.

[11] M. Krenn, M. Malik, R. Fickler, R. Lapkiewicz, and A. Zeilinger, “Automated Search for New Quantum Experiments,” Physical Review Letters 116, 090405 (2016). https://doi.org/10.1103/PhysRevLett.116.090405.

[12] C. Ruiz-Gonzalez, S. Arlt, J. Petermann, S. Sayyad, T. Jaouni, E. Karimi, N. Tischler, X. Gu, and M. Krenn, “Digital Discovery of 100 Diverse Quantum Experiments with PyTheus,” Quantum 7, 1204 (2023). https://doi.org/10.22331/q-2023-12-12-1204.

[13] K. Kasaura, K. Tsukamoto, K. Mori, R. Mizuno, T. Namatame, Y. Oriike, M. Taniguchi, S. Sonoda, and H. Yamasaki, “Lean-Quantum: Toward AI-Assisted Formalization of Quantum Information,” arXiv:2607.05492 (2026). https://arxiv.org/abs/2607.05492.

[14] L. Zhang et al., “Benchmarking Agents for Proving Theorems in Quantum Algorithms and Quantum Information,” arXiv:2607.21533 (2026). https://arxiv.org/abs/2607.21533.

[15] B. Alexeev, K. Barreto, Y. Li, J. D. Lichtman, L. Price, J. I. Shah, Q. Tang, and T. Tao, “Primitive sets and von Mangoldt chains: Erdos Problem #1196 and beyond,” arXiv:2605.00301 (2026). https://arxiv.org/abs/2605.00301.

[16] T. Tao, “Primitive sets and von Mangoldt chains: Erdos Problem #1196 and beyond,” What’s New, May 3, 2026. https://terrytao.wordpress.com/2026/05/03/primitive-sets-and-von-mangoldt-chains-erdos-problem-1196-and-beyond/.

[17] L. de Moura and S. Ullrich, “The Lean 4 theorem prover and programming language,” in Automated Deduction – CADE 28, LNCS 12699, 625–635 (2021). https://doi.org/10.1007/978-3-030-79876-5_37.

[18] K. Bharti, “Suppose a Lean proof confirms every logical step,” LinkedIn post, August 10, 2026. https://www.linkedin.com/posts/kishor-bharti-73824357_suppose-a-lean-proof-confirms-every-logical-activity-7492780443344216064-2AfO.

[19] K. Bharti, “QIP may eventually include more talks that explain a whole direction,” LinkedIn post and public discussion, August 11, 2026. https://www.linkedin.com/posts/kishor-bharti-73824357_qip-may-eventually-include-more-talks-that-activity-7493097099555790848-6Q65.

[20] K. Bharti, “AI can be very useful in peer review,” LinkedIn post, August 9, 2026. https://www.linkedin.com/posts/kishor-bharti-73824357_aican-be-very-useful-in-peer-review-it-activity-7492209471818059776-q2uc.

[21] K. Bharti, “Some disclosure of AI use is also probably useful,” LinkedIn post, August 10, 2026. https://www.linkedin.com/posts/kishor-bharti-73824357_some-disclosure-of-ai-use-is-also-probably-activity-7492635721451798528-4CFI.

[22] K. Bharti, “If an AI immediately gives the right counterexample or proof strategy,” LinkedIn post and public discussion, August 7, 2026. https://www.linkedin.com/posts/kishor-bharti-73824357_if-an-ai-immediately-gives-the-right-counterexample-activity-7491640325204164608-vkIG.

[23] K. Bharti, “AI can lower barriers in a very real way,” LinkedIn post, August 8, 2026. https://www.linkedin.com/posts/kishor-bharti-73824357_ai-can-lower-barriers-in-a-very-real-way-activity-7491831700155613184-NgVF.

[24] K. Bharti, “There is an easy pessimistic story in which machines prove theorems,” LinkedIn post and public discussion, August 12, 2026. https://www.linkedin.com/posts/kishor-bharti-73824357_there-is-an-easy-pessimistic-story-in-which-activity-7493383724651376640-jkuX.

[25] R. Joober, N. Schmitz, L. Annable, and P. Boksa, “Publication bias: what are the challenges and can they be overcome?” Journal of Psychiatry and Neuroscience 37, 149–152 (2012). https://doi.org/10.1503/jpn.120065.

[26] T. Haug, K. Bharti, and D. E. Koh, “Pseudorandom unitaries are neither real nor sparse nor noise-robust,” Quantum 9, 1759 (2025). https://doi.org/10.22331/q-2025-06-04-1759.

[27] K. Bharti, “Pseudorandom unitaries are neither real nor sparse nor noise-robust,” LinkedIn post, June 20, 2023. https://www.linkedin.com/posts/kishor-bharti-73824357_pseudorandom-unitaries-are-neither-real-nor-activity-7077106652646166529-_tMs.

[28] Quantum Information Processing Conference, “QIP Charter,” accessed August 13, 2026. https://qip.iaqi.org/qipcharter.

[29] S. Gharibian, “Social evolution of QIP talks moving forward,” LinkedIn post and public discussion, August 4, 2026. https://www.linkedin.com/posts/sevag-gharibian-a148aaa4_social-evolution-of-qip-talks-moving-forward-activity-7490417087774085120-9Ne6.

Comments

  1. Hi Kishor,

    This is an excellent overview of the mental dilemma researchers are grappling with. Esp. as someone working at the interface of QC and AI, and effectively working to automate quantum engineering, it was particularly enlightening!

    It's pretty long though, and I believe, (to quote Blaise Pascal), if you have had more time, you would have written a shorter post. :D

    I had a few thoughts as I was reading through, and wanted to pen them down.

    Some things I resonated well with:

    1) "..credit system built mainly around temporal priority. If several strong systems can search a space of ideas in parallel, being first may depend on access, budget, orchestration,.."

    2) "The durable contribution may lie elsewhere: who identified the right problem, who understood the proof, who found the clean abstraction, who checked the security definition, who connected the result to earlier work, and who made the method reusable." - research would indeed be measured more via accessible/useful/grounded abstractions and current/potential impact.

    3) "The human role does not remain safe merely because a laboratory contains hardware." - given a lot of current focus is on physical AI, lab automation would soon be feasible.

    4) "If we define human value as whatever today's system cannot do, we will be forced to retreat each time the capability boundary moves. The human role should instead be stated in terms of responsibility. People and institutions must decide which questions merit resources, which formalization preserves what matters, when evidence is strong enough to call a claim knowledge, and what consequences follow from acting on it."

    5) "We may respond to cheaper proof production by demanding more theorems per paper, more experiments per theorem, or a faster publication cycle. That would preserve the appearance of scarcity by raising the quantity threshold. It would also move us toward a literature that almost nobody can read. A healthier response is to value the work of selecting, compressing, and explaining."

    6) "That proposal is attractive, provided the record does not become a demand to publish every private conversation or every failed prompt. Research requires space for exploratory thought. Provenance should make a contribution legible, not turn intellectual work into surveillance."

    7) "No policy will work if honesty is punished. If authors believe disclosure lowers acceptance, hiring, or grant prospects, the most careful researchers will be more transparent and may be penalized for it, while others remain silent."

    8) "Scientific English has long acted as a gate. Reducing that gate is not cosmetic. A good idea that cannot be presented in the expected style is often not evaluated fairly. At the same time, a new inequality may appear. The advantage will not belong only to people who can afford a subscription."

    ReplyDelete
    Replies

    1. While some, I am still debating with myself:

      1) "If AI increases the supply of technically polished results, the program committee will not only be filtering correctness. It will be deciding what is worth asking hundreds of people to understand together." - and what prevents program committee members from using AI to filter through submissions, given it's pretty routine for HR to use ATS to shortlist interviewees profiles.

      2) "The work of checking, reconstructing, comparing with the literature, and making the idea intelligible remains real scientific labor." - not sure for how long this would remain a forte.

      3) "An author must answer questions, correct errors, disclose conflicts, respond to criticism, and bear reputational responsibility." - in principle nothing prevents AI models from doing this. You can automate a scientific career for Claude Fable 5, make it reply to errors and criticism, and audit citation on a Google scholar profile without a social security number.

      4) "A useful default is human first, AI second. The reviewer should read the abstract and introduction, inspect the main result, and form an independent view before asking a model for a summary or critique." - while I follow this and agree in principle, it is really tempting to automate the full reviewing pipeline to just prompt the AI for both the verdict and the review with the context of the full venue and the article, esp. when not all reviewers have an aligned moral compass and firefighting deadlines. I am really trying to shrug off Yudkowsky's AI-fatalist ideas from creeping in, but honestly, the human territorial dominance is fast shrinking.

      5) "... has publicly suggested short videos as evidence of how the humans responsible for a result communicate it, regardless of how it was generated." - I doubt it that would qualify as responsibility accounting when it is so easy to generate summary scripts, video from single profile photo, audio mimicking our voice, and accompanying explainer slides.

      I think some of my writeups might interest you: https://aritrasarkar.com/musings/agi/

      Delete
    2. Aritra, thank you for reading this so carefully and for the thoughtful comments. And yes, the Pascal criticism is fair :)

      I agree with much of what you raise. I do not think reviewing, verification, explanation, or responding to criticism are stable “human territories.” AI may become excellent at all of them.

      The distinction I was trying to make is less about capability and more about responsibility. Even if AI performs most of the intellectual work, someone still has to decide what to trust, publish, fund, or act on, and be accountable for those decisions.

      Your point about reviewing is especially important. “Human first, AI second” may well be only a transitional norm. My concern is mainly that we should not silently let an automated filter define what the community considers interesting or important.

      More broadly, I think trying to define the human role by what AI cannot do is probably a losing strategy. The better question is what kind of scientific community we want when AI can participate in almost every part of research.

      Thanks again. These are exactly the kinds of objections I was hoping the post would provoke. I will also read your AGI writeups.

      Delete

Post a Comment