Debating an AI Opponent

I have written on several occasions about the educational potential of debate. Researchers who focus on this topic often use the alternative term “argumentation“. Debate requires a deep understanding of the contested topic and offers a concrete way to exercise critical thinking. It also involves a motivational element because it involves competition. I read some educational research on debate and made a personal connection, noting how quality argumentation differed from the arguments I read on social media. Most of my previous posts about argumentation are available on my blog

A couple of years ago, I encountered a proposal that AI chatbots could serve as debate partners and saw a connection to Deanna Kuhn’s research, which investigated text messages as a way to develop argumentation skills because the process created a written record that could be analyzed and discussed. I saw a similar opportunity in the written record produced in an AI chat and the AI system could provide an opponent. 

I had done little more with this idea until recently, but I began thinking about a more formal debate format and decided to see how AI would implement the format I had in mind. 

Here is the prompt I have been exploring.

I want you to debate me on the topic of whether a computer and keyboard or handwriting on paper is the best note-taking approach. First, write a 200-word position statement supporting handwriting, and I will do the same for using a keyboard. After evaluating my position, write a challenge to my position. I will reply and challenge your position.

I chose this topic because I have read much of the related research and popular press and I think I understand the issues and competing positions. I also argue the less popular digital approach. This combination seemed a good combination to evaluate the experience. 

The change I made from my previous approach to an AI debate was to add the expectation that each participant create an introductory position the two parties a way to identify what appeared to be issues that were in disagreement. One more thing, while I first explored this experience using an AI subscription I use (Claude), I then tried the Mac mini on which I self-host an AI system (running Ollama) and repeated the process. The self-hosted experience was great and had the strange benefit of being to take an hour or so off for lunch and come back to the computer with the debate still active. I was particularly impressed by the content the AI was able to through at me. This Mac was not connected to the Internet and the only explanation I can offer is that I use the same setup to run a Karpathy Wiki and I have a great deal of relevant content within that system.

I have a substantial amount of content stored from this session and I often include such material as an Appenedix, but this was simply too much material. As an alternative, I decided I would upload the transcript of the entire debate to NotebookLM (Gemini Notebook) and include a link to this notebook for anyone interested in reading the transcript. 

I had one final idea. What if I use the AI embedded with the notebook to ask for an evaluation of the arguments and a winner?

Even if you don’t examine the transcript of the debate, NotebookLM’s evaluation gives you a sense of the level of the content and the interaction produced. I was a little disappointed because I thought my position was beyond logical refutation, but Notebook seemed unwilling to declare me the clear winner. The experience was not unlike actual debates. Can you get your opponent to engage with you on what you want the judge to accept as the key or deciding issues? 

Prompt: I want you to evaluate this debate considering whether taking notes by hand or using a computer is best in preparing a learner for tests based on lecture content. Are the best arguments made by the pro-keyboard or pro-handwriting supporter?

Evaluating this debate through the lens of preparing for tests based on lecture content reveals a fundamental tension between efficient information storage and active cognitive processing.

Both sides offer strong, distinct arguments depending on the type of test being taken and the student’s studying habits.

The Pro-Handwriting Case: Best for Conceptual Tests and Internal Recall

The pro-handwriting supporter makes the most compelling arguments regarding internalization, conceptual understanding, and performance under actual exam conditions.

  • The “Recall” vs. “Retrieval” Advantage: In an exam setting, students do not have access to their devices or a search engine. The handwriting supporter argues that digital note-taking promotes “cognitive offloading” and “retrieval” (knowing where to find a fact in a database) rather than “recall” (pulling a concept from internal memory). Handwriting forces the brain to do the heavy lifting immediately, creating the robust neural pathways needed to recall information under pressure without a digital “crutch”.
  • Performance on Conceptual Questions: Backed by the Mueller and Oppenheimer (2014) study, the handwriting advocate highlights that because typing is fast, it leads to mindless “transcription”. Because physical writing is slower, hand-writers must summarize and prioritize key concepts in real-time. This “selective encoding” means they perform significantly better on conceptual and synthesis test questions.
  • The “Sieve” vs. “Bucket” Study Cycle: Handwriting acts as a “sieve” that filters information during the lecture. When test preparation begins weeks or months later, the hand-writer reviews a distilled, pre-processed “mental map”. In contrast, a typed “bucket” of notes requires the student to spend limited study time filtering a mountain of raw, flat data before they can even begin deep synthesis.

The Pro-Keyboard Case: Best for Detailed Tests and Long-Term Study Flexibility

The pro-keyboard supporter makes highly sophisticated arguments regarding information completeness, managing cognitive load, and the reality of long-term study cycles.

  • Performance on Detail-Oriented Questions: The keyboard allows students to type fast enough to capture almost everything accurately. For exams that test factual details and specific data points, the pro-keyboard side is heavily favored, as research indicates laptop users perform better on detail-oriented questions due to their comprehensive records.
  • Managing Lecture Cognitive Overload: The keyboard supporter rightly points out that a fast-paced lecture is not the ideal setting for deep conceptual thinking. Attempting to synthesize while keeping up with a speaker can cause cognitive overload. Keyboarding minimizes the friction of capture. This “offloads” the processing during class so the student can focus on capturing data, leaving the deep reflection to a low-pressure study environment.
  • Iterative Study and Optimization: Digital notes are highly malleable. Features like searchability, audio-linking, and inserting placeholders (like “XXX”) for missed content allow students to efficiently patch gaps, reorganize ideas, and elaborate on notes over a semester-long period. This active post-processing provides a substantial generative advantage over rigid, handwritten pages that are cumbersome to restructure.

Conclusion: Who Makes the Best Arguments?

The pro-handwriting supporter makes the superior psychological and cognitive arguments for final test performance. Their points are deeply anchored in established learning science—such as “desirable difficulty” and the “generation effect”—proving that the physical friction of writing is a feature that directly constructs a student’s internal cognitive map. Under test conditions, a “remembered concept” will always outperform a “searchable note”.

However, the pro-keyboard supporter makes the more realistic argument regarding modern study workflows. If a student is highly disciplined and actively uses the study phase to review, edit, and link their digital notes, the digital “workspace” provides a highly scalable and superior infrastructure for long-term learning.

Ultimately, if the goal of note-taking is to build a high-functioning cognitive map to ace conceptual exams, the pro-handwriting arguments are the strongest. If the goal is to build a durable, searchable archive for detailed reference, the pro-keyboard arguments win out.

Would you like me to create an interactive study guide or a set of flashcards based on these opposing cognitive theories to help you further analyze these learning principles?

References:

Kuhn, D., Goh, W., Iordanou, K., & Shaenfield, D. (2008). Arguing on the Computer: A Microgenetic Study of Developing Argument Skills in a Computer-Supported Environment. Child Development, 79(5), 1310-1328

Mayweg-Paus, E., Macagno, F., & Kuhn, D. (2016). Developing Argumentation Strategies in Electronic Dialogs: Is Modeling Effective. Discourse Processes, 53(4), 280–297. https://doi.org/10.1080/0163853X.2015.1040323

Loading

AI (Grammarly) and Writing: Good or Evil?

I am interested in the potential of AI in developing writing skills. I am presently focused on Grammarly, having used the tool for years, and am now considering how it might play a productive role in secondary and higher education efforts to develop writing skills. Exploring what those writing about Grammarly on Medium have to say about Grammarly, I have come away with the impression the majority argue the tool is overpriced, less effective than tools with a similar purpose, and generally a bad idea when applied in an educational setting. I don’t agree. 

Writing in the classroom

I like to draw a distinction between learning to write and writing to learn. This distinction is artificial, as classroom instructional strategies such as “Writing Across the Curriculum” argue that both goals can be addressed when writing assignments in other disciplines are evaluated both for the quality of the writing and for what the writing suggests about students’ understanding of a given topic. My argument here focuses on the potential value of AI in learning to write. 

When and why is AI a problem when learning to write

In situations where the development of writing skills is the emphasis, an AI tool is argued to be problematic because students cheat by using AI to avoid doing work that requires them to practice the skills they are expected to master. In addition, by turning in work for evaluation that students did not actually perform, they do not receive feedback on the skills they are supposed to be learning and are credited with achievements they have not actually earned. 

AI and writing: A different take on the actual problem?

Many educators, aware of the possibility of cheating, have resorted to approaches such as short, handwritten in-class assignments that eliminate the possibility of using AI. There are limitations to this approach, especially for the unique skills required to create longer arguments or other lengthier projects. 

As adults with reasons to write and without worrying about the need to prove everything that appears in a written product is based on our own knowledge and writing skills, we may take advantage of AI in many different ways. One real question is how, and perhaps if, we are preparing students to transition from a focus on learning new skills or the graded demonstration of one’s knowledge to a combination of AI and personal knowledge and writing skills. In addition, some are suggesting, and again rightfully so, that AI can benefit students’ efforts to learn to write and write to learn. Here, I want to emphasize the potential assistance in learning to write more effectively. 

I tend to react to what I think are naive expectations of teachers and the reality of working in classrooms is important here. For example, I support the exploration of AI as a tutor, not because I think AI is equivalent to a human tutor, but because human tutoring is costly and many students who need help do not receive sufficient human attention as a consequence. I have a similar opinion about learning to write. It would be great if each student could write a lot and receive rapid feedback, as well as an individual conference related to their effort. Neither immediate, consistent feedback nor frequent individual attention is practical. Just having an AI writing tool, such as Grammarly, that can provide immediate feedback on what has been written seems like a practical improvement.

So Much Depends on Personal Motivation

Grammarly and asking pretty much any AI tool to evaluate specific attributes of your writing quality will provide you with feedback to consider. The issue is really whether you take the time to ask for this feedback and to consider the feedback that is produced. Here is what I mean. I use Grammarly while I write, and it constantly provides feedback. In reflecting on my own behavior, I almost always quickly accept the suggestions for what I have written (these appear as underlines in various colors) by clicking to have Grammarly fix the problem. I don’t stop to figure out what was wrong with what I wrote. Was that an actual error of grammar or spelling, and if so, why? The fixes always seem better, but they also remove what may just be my voice or personal preference in how I say something. I avoid the opportunity to learn and also allow Grammarly to “standardize” my writing. At this moment, admitting this has made me self-conscious. 

This reminds me of the experience I had providing comments on many of my grad students’ theses and dissertations. In later years, I liked to use the comments feature in Google Docs to leave comments and identify actual errors. I started to realize that some students were simply allowing me to rewrite their papers, when what I wanted was for them to consider something different. Often, I had to remind them of the difference between my thoughts about their work and the actual errors I pointed out. 

If you use a tool such as Grammarly, you probably recognize my observation in your own behavior. It is so easy to accept proposed changes based on a kind of “that sounds pretty good thinking” and trust in the assumption that the “system knows the rules better than I.” Taking this approach is quick, painless, and “good enough.” The problem is that this approach fails to take advantage of at least some of these situations to learn. Why were these changes recommended? Is my way of expressing myself flawed or just unique? Grammarly will help you consider which is most likely. 

What was wrong with what I wrote?

Grammarly has always allowed you to pause when suggesting a change. There was no time limit on the opportunity to consider what you wrote in comparison to what was recommended. As the tool was improved and with the more recent integration of AI, efforts were made to explain why a change was recommended. At first, the tool offered a general reference to rules. Here is what a split infinitive is, and here are some examples of sentences containing a split infinitive and improved versions of the same sentences. Here is an example of passive voice, and here are some examples. The most recent advance offers similar information, but specifically related to your own words rather than just generic examples. 

One note – I have encountered descriptions of this newest capability from others, but I haven’t been able to replicate the same output on my own computer with the latest version of Grammarly. I have had this difficulty even though I input exactly the same text used in the other demonstrations I have encountered. My setup will identify the error and provide generic examples, but it won’t explain based on the text I have entered. I can generate explanations specific to my written text, but I have to use the AI window to enter a prompt asking for this information (see examples below). 

Here are a couple of examples. In the first, you see a sentence with three components underlined in blue (I highlighted it in blue so you can find it). In the associated column on the right you see the proposed alternative with the changed words or punctuation bolded. The red box identifies the button to get additional information. The second image shows the result of making use of this button. The explanation for the proposed changes – Our revised version of your text may make it more effective. – is Grammarly’s explanation. As I said, this offers little of value to me. However, if I use the AI prompt box (Grammarly calls this Superhuman Go), which appears at the base of the right-hand column, I can ask for an explanation, and I receive the lengthy explanation shown in the third image. 

Here is a second example that explains a split infinitive. Again, if I ask for an explanation related to the limitation in my writing, I get both a general and a targeted explanation of the split infinitive (second image).

Improvements at multiple levels

Improvement in writing happens at multiple levels. There are spelling errors and grammatical failures. These are the easiest for a digital tool to identify, and spelling correction has been around for decades. These might be described as superficial errors. There are deeper levels of good writing that any of us who critique others’ work find much harder to evaluate and that are the most difficult proficiencies to teach and learn. Grammarly’s AI capabilities provide a way to address these proficiencies. The easiest way I can describe this capability is to list some of the prompts Grammarly suggests. 

  • “Critique the logic and reasoning in this text.” 
  • “Are there any gaps in my argument here?” 
  • “Does this conclusion follow from my premises?” 
  • “Point out any weak points in my reasoning.” 

I will append Grammarly’s extensive analysis of a recent post based on one of these prompts (see Appendix). I have extensive experience as an academic writer with two published textbooks. These works involved feedback from a dedicated editor and input from multiple paid professors regarding both the content and communication style. The detail in Grammarly’s feedback is easily as detailed. 

Summary

The point I am trying to make here is not that AI cannot be used to cheat. My point is that Grammarly and other AI tools can also provide feedback more quickly and consistently than is practical for educators. This seems to be another example of the constant dilemma with AI, it can be used for good or for evil. In this case, when a writer is working on their own, they decide which it will be. 

Addendum

Grammarly’s analysis of AI Processing Options

Prompt: Point out any weak points in my reasoning

A few places in the reasoning could be stronger or more precise:

  1. The continuum needs clearer criteria
    You describe a continuum from “total AI” to “total human,” but the dimensions of that continuum are a little mixed. Sometimes the distinction is about who selects the source material, sometimes who reads it, sometimes who organizes it, and sometimes who writes the final text. You might strengthen the piece by naming the dimensions explicitly, such as:
  • control over source selection
  • depth of human engagement with sources
  • human contribution to interpretation
  • AI role in organization
  • AI role in final wording
  1. “Information input” may be too narrow
    In the Recall.ai example, you say the main thing you control is the information input. But choosing 15 PDFs is not just supplying information; it also frames the argument by deciding what evidence counts. That choice may involve substantial judgment, even if you do not read the documents closely.
  2. The quality question needs more defined standards
    You ask which approach produces the highest-quality product, but “quality” could mean many things: factual accuracy, originality, usefulness, writing style, source integration, depth of understanding, efficiency, or ethical acceptability. The argument would be stronger if you noted that different approaches might score differently depending on the quality standard.
  3. The “total AI” end may not be truly total AI
    Even in the ChatGPT/Claude example, the human still defines the topic, purpose, and prompt. So it may be more accurate to call this “minimal human input” rather than “total AI.”
  4. The “total human” end still includes tool-mediated thinking
    The Obsidian-only example is described as having no AI use, which makes sense if the point is AI involvement. But it still involves a digital tool that shapes organization, retrieval, linking, and thinking. If your continuum is specifically about AI involvement, that is fine—but you may want to clarify that it is not a continuum from tool-free writing to machine-written writing.
  5. The RAG example may understate verification
    You note that you do not have to read the documents, but that raises an important issue: without reading them, you may not be able to evaluate whether the AI’s summaries are accurate or whether the final product fairly represents the sources. This is an important weakness in that approach and could be made more central.
  6. The Karpathy Wiki stage could use more explanation
    This is the most complex example, but the reasoning around why it is distinct could be clearer. Is the key difference that AI creates a reusable knowledge structure? That it separates source processing from final generation? That it mimics a human note system? Stating the main conceptual difference directly would help.
  7. “Parallel to how human writers function” is promising but underdeveloped
    You suggest that identifying themes and concepts before drafting may parallel human writing. That is an interesting claim, but it needs a bit more support. Human writers do not just extract concepts; they evaluate, synthesize, prioritize, and sometimes change direction. You could strengthen this by explaining exactly which parts of the AI process resemble human writing and which do not.
  8. The ethical/authorship issue is implied but not addressed
    Your categories raise questions about authorship, responsibility, and acceptable AI assistance, but those issues are not directly discussed. If this is meant to help readers evaluate AI writing practices, you may want to include a brief note about where responsibility for claims and interpretations resides.
  9. The examples could be labeled more consistently
    Some category labels name tools, while others name processes. For example, “ChatGPT or Claude” names tools, while “Obsidian to Karpathy Wiki to Product” names a workflow. You might make the structure more analytical by using labels like:
  • Prompt-only generation
  • Source-constrained generation
  • Human-annotated source generation
  • AI-organized knowledge-base generation
  • Human-authored synthesis

Overall, the central idea is strong: you are not treating “AI writing” as one thing, but as a range of workflows with different levels of human involvement. The main improvement would be to make the continuum’s criteria explicit so readers can see exactly what changes from one category to the next.

Loading

An AI LLM Can Identify Important Missing Content In Student Notes

There seems to be a resurgence of research interest in note-taking and note-use. One source of new work was the controversy over whether students should take handwritten notes or use their laptops or tablets. A second focus has been generated by what I would describe as “out of classroom” notetaking described by those promoting personal knowledge management (PKM) and interactions with a second brain. 

I have had a long-term interest in notetaking and studying and recent activity has caused me to review old sources in light of newer activity. Possible connections between notes taken and stored digitally and the capabilities of AI tools seem worth exploring.

I still understand note-taking and note use from a decade old model that argues learning with and from notes can be outlined in terms of an immediate generative benefit of taking notes and by then using the external storage function of notes for review and study. A generative benefit implies that taking notes provides an advantage over simply listening and reading. The external task prompts the note-taker to identify important ideas because it is not practical to record everything, and then cognitively acts on the selected information to personalize the original content through summarization and connections to existing knowledge. Of course, there are many strategies for doing this, as well as individual differences and less active approaches, such as simply writing down what the instructor displays on his/her PowerPoint slides.

One topic relevant to how I imagine AI being used in interaction with student notes is work that examines how effectively learners have recorded important ideas. How many ideas find their way into the notes and are these ideas important ideas? Here you find some interesting controversies. Early work on note-taking in actual classroom settings found a substantial correlation between the volume of content students recorded and their exam performance (Nye and colleagues, 1984). I want to emphasize the word “naturalistic” here because, as the authors of this study claim, so much of the research has been conducted in short-term laboratory studies involving short segments of content and immediate examinations – not the situation students face in practice. I have made the same complaint in a recent post

A recent controversy related to what might be described as the “more is better” position is the handwriting vs. keyboarding controversy. The issue here is that keyboarding allows for more input, but handwriting appears to lead to more generative activity because the slower process encourages greater selectivity and summarization. My observation is that selection and reworking of notes does not have to happen in real time and having more content to work with during review and study makes such reworking more productive. 

This brings me to a second observation. While it seems possible that learners would not record information they think they already know, for one reason or another, students fail to record a sizeable proportion of important information, and what is not recorded can be related to specific later failures of recall and application on learning measures. It is difficult to study what you don’t know and what was not recorded (e.g., Kiewra, 1985; Lawson & Mayer, 2024).

Expert Notes and the AI Alternative

One solution to the many important ideas that students miss, especially less capable students, is to provide what are often called “expert notes” – essentially notes provided by the instructor or taken by an advanced student (Kiewra, 1985). This is a topic I explored in my own research with my graduate students (e.g., Grabe and colleagues, 2005). To be clear, the idea is not to substitute expert notes for student notes because of the generative value of taking notes, but to provide a second source learners can use to “fill in the blanks” for what they missed or perhaps did not understand well enough to even attempt recording. 

There are practical issues with providing quality notes to students. One issue that interested me in the citation I provided was the relationship between providing quality notes and class attendance. You do want students to attend class, but the reality is that there is also the question of what can be done for those students who miss class for legitimate reasons. Second, the generative effect of taking notes describes note-taking itself as a learning experience, so you really don’t want to discourage student note-taking. 

I have been exploring a use of AI that seems interesting and efficient in a couple of ways. It is relatively easy to prompt AI to identify the “important ideas in a lecture”, to compare this material with the notes taken by a learner, and to generate an output of the important ideas not contained in the student notes. This approach is efficient for students because it does not require reviewing the entirety of what might be considered “expert notes.” It also builds on rather than eliminates the student’s generative effort to create their own notes.

To explore how this might work, I used my current AI service to run the following prompt. The inputs included, first, the transcript of a source podcast on why climate change can be responsible for both drought and flooding, and, second, the notes I took while listening to the podcast. The original source contained approximately 2500 words and as a podcast required approximately 15 minutes. I attempted to play the role of students and created a set of notes while listening. I admit that I was surprised to learn a few things about how climate change works even though the material was prepared at the high school level. I purposely ignored comments that appeared near the end of the podcast that focused on fire danger and health concerns as a way to approximate what notes might look like if a student became distracted or cut out of class early. 

Prompt: I am providing two text segments. I want you to identify the important content from the first segment. I then want you to compare this important content with the content in the second segment and identify the important ideas in the first segment that are not included in the second.

I have included, as an Appendix, both the important ideas identified by the AI tool and the list of “important ideas” in the presentation that the AI tool said were not included in my notes. 

The approach identified and provided missing comments on both wildfires and health issues, which was my proof of concept when it came to responding to the problem of students missing important ideas, but it also included additional information that the comparison of the original content and my notes did not. So, when I compared the summary with my notes, the comparison in response to the prompt identified content I purposefully missed and other items I heard but did not record. This second body of information was much larger than I would have anticipated and fits existing research concluding even better students (apologies if this sounds conceited) fail to record more than one might expect. It is hard to know if this is an actual problem or simply a difference in standards. I have no idea whether the prompt could be altered to change how the AI tool interpreted “important content.” 

One interesting item – there is at least one item in the comparison between the AI summary of important ideas and my notes identified as missing that actually does not appear in the AI summary but was included in the full transcript (The claim that the most destructive floods now occur about three times as often as before). 

Final Comments

It is likely not a surprise that AI creates a summary of content, as this is a commonly offered service now often available as an alternative to asking readers or listeners to review all of the original raw content. I am not certain if I explored different wordings of my prompts if a request for the “important ideas” would be much different from a summary.

I have considered whether a summary or review of the most important information would offer essentially the same practical value as the second step I have AI to provide. The amount of content would differ. The summary I provide as an appendix is about a third of what would be generated from a traditional fifty-minute college lecture. I do like the idea of identifying what a student missed, which may have unique value in prompting students to consider their note-taking process. Helping students identify important information they missed also keeps the focus on student responsibility for the generative learning opportunity in making their own notes.

Sources:

Grabe, M., Christopherson, K., & Douglas, J. (2005). Providing Introductory Psychology Students Access to Online Lecture Notes: The Relationship of Note Use to Performance and Class Attendance. Journal of Educational Technology Systems, 33(3), 295-308.

Kiewra, K.A. (1985). Providing the instructor’s notes: an effective addition to student notetaking. Educational Psychologist, 20(1), 33-39. https://doi.org/10.1207/s15326985ep2001_5

Nye, P.A., Crooks, T.J., Powley, M., & Tripp, G. (1984). Student note-taking related to university examination performance. Higher Education, 13(1), 85-97. https://doi.org/10.1007/BF00136532

Appendix:

Important content in Segment 1

  • Climate change is intensifying both droughts and heavy rainfall/flooding.
  • A warmer atmosphere holds more water vapor, while hotter land surfaces evaporate moisture faster:
    • This dries land and worsens drought.
    • It also creates the potential for more intense downpours once rain occurs.
  • The most serious change may be in extreme storms, rather than in average rainfall or average storms.
  • Existing infrastructure is poorly suited to extreme weather, especially because roads, roofs, parking lots, and other impervious surfaces increase runoff and flooding.
  • Flood risk is not limited to coasts and rivers; intense rainfall can cause serious stormwater flooding inland.
  • Communities can reduce flood impacts through:
    • Updated emergency plans and warning systems
    • Recovery resources
    • Green space, trees, grass, and rain gardens
    • Reducing impervious surfaces, including through greater public transportation
  • Climate impacts vary by region:
    • New England’s warming coastal ocean can fuel stronger winter storms.
    • Subtropical regions, including the U.S. Southwest, are generally expected to become drier or more arid.
    • Some areas may not follow this pattern; for example, the Sahara may become wetter.
    • Monsoon regions may receive rain in fewer, more intense events, harming people and agriculture.
  • Water-cycle changes also worsen other hazards:
    • Drought can increase wildfire risk.
    • Alternating wet and dry periods can grow vegetation and then dry it into wildfire fuel.
    • Humidity makes heat waves more dangerous because it prevents sweat from cooling the body effectively.
  • Adaptation is necessary, but reducing fossil-fuel use and climate pollution is also essential to limit future worsening.
  • Community discussion and engagement are presented as meaningful individual responses to climate change.

Important ideas from Segment 1 not included in Segment 2 (my own notes)

  1. Examples and evidence of current extremes
    • California reservoir lows, Fort Lauderdale flooding, and Brooklyn’s extreme rainfall.
    • The claim that the most destructive floods now occur about three times as often as before.
  2. The distinction between average conditions and extremes
    • The largest storms are becoming stronger even if the average storm is not necessarily changing as much.
    • The primary concern is extreme downpours, not merely average rainfall.
  3. Detailed flooding and infrastructure discussion
    • Impervious surfaces such as roofs, roads, sidewalks, and parking lots worsen runoff.
    • Inland communities face stormwater-flooding risks even when they are far from rivers or coasts.
    • Specific adaptation measures: rain gardens, trees, green space, early-warning systems, emergency planning, recovery resources, and public transportation.
  4. New England’s regional climate risks
    • Rapid warming of nearby ocean waters.
    • Greater land–ocean temperature contrast in winter that can help fuel nor’easters, blizzards, and hurricane-force winds.
  5. Regional variation and uncertainty
    • Climate impacts do not follow identical rules everywhere because continents and local wind systems complicate global patterns.
    • The Sahara may become wetter despite the broader tendency for many subtropical dry regions to dry further.
    • Monsoon rainfall may become more variable and concentrated into fewer, more intense events.
  6. Wildfire impacts
    • Drier conditions make vegetation easier to ignite and allow fires to spread.
    • Wet-to-dry cycles can build vegetation and then convert it into highly flammable fuel.
  7. Humidity and heat-wave danger
    • Humid air makes sweating less effective, making humid heat more dangerous than dry heat.
  8. Mitigation and social response
    • The need to transition away from fossil fuels and reduce climate pollution.
    • The message that some worsening is already expected, but further escalation can still be limited.
    • Encouragement for people to discuss climate change within their communities to build understanding and counter despair.

Loading

When is the Karpathy Wiki Better than RAG

This is a follow-up to an earlier post that concerned my attempts to implement the Karpathy Wiki strategy based on raw notes in Obsidian using local AI. Since that post, I have tried several AI models and ran the complete ingestion and wiki setup based on approximately 150 notes to completion without the weird inclusions I described in my initial post (Note: I will describe what I think was responsible for the unwanted inclusions at the end of this post.) The process required my Mac mini to run for approximately 26 hours. I have now reached the point I have opinions about whether the continually improving Karpathy wiki can deliver the type of insights my writing goals require.

I did not make use of my entire Obsidian note archive, and for the local setup I have, this appears necessary. This insight is consistent with other advice I have read concerning local setups, which recommend creating multiple Obsidian vaults if a user wants to address multiple categories of notes.

The focus of the Karpathy application I have been working with concerns notetaking approaches and issues. To reach conclusions regarding whether a wiki generated from my notes would satisfy my writing goals, I focused on the topic of whether handwritten notes are superior to notes taken on a device. I don’t know exactly, but I would guess maybe 30 of my notes would be relevant to issues associated with this topic. I have highlighted and annotated quite a few journal articles on this topic so what I describe as a note could contain several pages of material. I selected this topic to evaluate the Karpathy model because research on this topic seems to be inconsistent and I typically try to make arguments to explain the differences based on my analysis of the methodology of the studies. So, for example, in a vault containing many articles about note-taking, I might want to first identify those articles focused on keyboarding vs. handwriting, summarize the results, and, because I know differences exist in the way the researchers have reached conclusions about their results, investigate nuances in the research methodologies that might be responsible. It is likely that second brain approaches must vary with the eventual goals users are likely to pursue. My interests require relatively deep analyses rather than a focus on a shallower topic such as what is the general consensus of those who conduct research in this area. I did include annotations on the articles included in the exported notes, and it turns out that an important issue is whether a sense of my questions about article conclusions are reflected in the wiki content. You don’t really know when building a second brain what goals you may want your second brain to help you address a couple of years later. This seems a critical consideration when committing to the activity. Even if hints of where to look if different goals emerge, one would hope there would be a way to at least find the relevant source material stored in the second brain.

I can see several options. I could use my tags to find note-taking studies that involve handwriting and keyboarding, then reread my notes and possible sections of the original articles to add tags and perhaps generate a written study. Obsidian alone would provide this opportunity and a wiki summary would not be necessary. I could use a tool that has an embedded AI capability (say Mem.Ai, NotebookLM, or Recall.ai) to query the same collection of notes to see if I could find an efficient way to accomplish my task. I could apply an AI tool such as Claude, which I have been using in my Obsidian note system as a plugin and pay the token charges associated with the Claude API. I could create a Karpathy-type wiki to preprocess my notes and then search the wiki using AI (in my case, Ollama so I can run the AI locally) because this approach is being recommended as a way to avoid the typical RAG approach of ingesting large collections of notes with each new investigation. So, it seems there are categories of approaches that exist on three levels.

I understand that any effort to compare these approaches that may be labeled as note organization and preprocessing, RAG AI applied to the raw notes, and AI applied to the output of a Kaparthy wiki AI preprocessing is oversimplified and ignores differences in the cost, sophistication, and design of the different systems. However, on a personal level requiring me to use my own finances and relevant knowledge, I have at least differentiated several approaches, applied the same approaches to the same original content using the same information goals each multiple times, and generated some impressions. These are the impressions I offer here.

The issue with RAG seems to be cost and time to ingest the content the system users want to query. If I am committed to an extended examination of content to achieve a purpose — say, write a blog post I might ingest designated materials and then apply multiple prompts within the same session to find the approach and the output that allows me to achieve my goal. I have found this to be efficient and the costs manageable because the large number of tokens to ingest is expended once and the much smaller number of tokens to query several times. When I want an approach that may involve addressing many different goals over an extended period of time, the multiple ingest costs focused on the same content encouraged my interest in developing a “wiki” as a permanent intermediate step in getting from my original content to working to achieve my goals.

I admit that I don’t know the details of how the category of note use I have described as intermediary works. Some results of processing appear to be persistent. I upload my notes once and then can prompt these notes repeatedly multiple times per month for the same cost. I know these services save my previous uses of these notes because I can review previous prompts and results. My differentiation of this category from the type of RAG system a Karpathy wiki argues is inefficient is based mostly on the argument that a wiki solution is a useful innovation. It is unclear to me why the intermediate approach I describe is not preferable to what Karpathy proposes. I have to leave this issue for the time being. My point here is to identify what I see as the existing categories. I certainly welcome comments if readers can provide additional clarity or a correction to my speculation.

The Karpathy Wiki

The several scripts I have used to guide the creation of the Karpathy wiki have all generated three categories of documents: sources, entities, and concepts. I have included an example of each below, using examples that I hope you will see as related.

The sources are what I would describe as abstractions or summaries of the input files. Source is an unfortunate label, in my opinion, because it is not the original; it is based on the input and created by the AI.

An Entity is a specific “thing” or “noun” — a discrete object with a name. In my situation, these might be the names of the researchers who wrote the original articles I read, their institutions, or perhaps the name of an online service or a technique I included in the input I provided.

Concept is an abstract idea, a theory, or a methodology. It is interesting what the AI identifies as a concept. For example, “note-taking” and “note-making” are included as concepts because the terms appear repeatedly. Computer notes and handwritten notes are also listed.

These components are all cross-referenced within the wiki, and the links accumulate as additional raw inputs are processed. For example, an author (entity) may be connected to multiple sources, each based on a different journal article that the individual worked on. A concept may be connected to the work of several authors, or to those prominently mentioned in several inputs, and to different concepts mentioned in individual or multiple sources.

Source

Concept

Entity

Reactions

I have tried to select examples from each of the wiki categories that fit the goal I am using as an example. You can see that “hand-written vs. keyboarding” appears as a concept and relevant studies and researchers were identified from the input materials to appear in the wiki. There are mentions of differences in methodology associated with areas I believe are important (e.g., what can be gleaned from stored notes over time). Given that I have an existing opinion about this body of research and have an opinion on how the reported results may be misleading, I can find elements in the wiki I could use at least to get back to the more detailed original notes and to the full original articles. A more important issue for me is whether this would have been the case should investigating this possibility have been a new interest and not a perspective I had when creating the original notes.

Finally, exactly how the wiki is to be used seems to be based on a different approach. I understand Karpathy to focus on applying AI queries to the wiki content so you don’t have to take the typical RAG strategy of ingesting all of the raw material each time. I have included the return to two AI queries asking the same question — one using the RAG approach and the second using the wiki. I have decided that both would have given me a starting point for investigating why the results of research on the topic of interest varied, but I would likely have used the wiki-based results to prompt a RAG query because of the much greater detail of the RAG approach. More relevant research was identified than was found in the wiki approach and the analysis in the RAG-based approach is far more detailed. Across many writing tasks, this combination would have been efficient. Again, the two outputs I am asking you to compare both started with exactly the same collection of notes. These two prompt responses have been included as appendices.

Final comments on setting up the Karpathy wiki using Obsidian and Ollama on my Mac mini.

I identified the issues that resulted in the unintended content described in my initial description of my experience. These issues were fixed by clean installs of both Ollama and Obsidian (with plugins). Both Ollama and Obsidian generate JSON files that, once established, influence future efforts. In my initial explorations and false starts, I created issues that required a clean start.

The second issue is more embarrassing. Installing Ollama added a way to set specific features I did not notice that appeared at the top of the Mac screen. Settings I did not notice were easily identified when taking advantage of what this option displayed. I now believe it is possible that the default setting allowing access to the Internet provided a way for Ollama to find concepts and entities related to the notes I knew were not from my notes. Anyway, this was easy to block and things worked perfectly when I got rid of the JSON files and changed this setting.

Appendices Output based on the Karpathy Wiki

The research comparing handwritten notes with those taken on computers has yielded varying results, which highlight different aspects of cognitive processing and learning outcomes.

Handwritten Notes

  • Benefits: According to the journal summary The effects of writing on learning in science, social studies, and mathematics: A meta-analysis (citation [[WIKI_FOLDER/sources/notesvextended_e62af5|The effects of writing on LSSM learning]), handwriting notes can enhance conceptual understanding and retention. This is because the act of handwriting involves a higher level of cognitive processing which aids in deeper engagement with material.
  • Empirical Support: A study by Luo et al., highlighted in Repetition of Input and Notes (citation [[WIKI_FOLDER/sources/repetition-of-input-and-notes_7e2f6d|Repetition of Input and Notes]), suggests that repeated read-throughs enhance recall when notes are reviewed, suggesting the benefits of detailed note-taking.

Computer-Notes

  • Benefits: According to Repetition of Input and Notes (citation [[WIKI_FOLDER/sources/repetition-of-input-and-notes_7e2f6d|Repetition of Input and Notes]), typing notes can be more transcription-oriented, which may not encourage the same level of cognitive processing as handwriting.
  • Challenges: A key study by Luo et al. (2018) found that longhand note-taking was more generative during review compared to laptop notes, which tended to be more superficial and less effective in fostering deep learning (citation AI response to my question challenging the Luo study that assumed the notes themselves could not result in generative activity).

Note Modification

In Conclusion

The evidence indicates that while handwriting may offer advantages in deep cognitive engagement and memory retention, typing notes on a computer can equally enhance learning outcomes if supported by effective note-taking strategies and subsequent revisions. The key difference lies in the type of engagement required during the initial note-taking process-handwriting encourages more direct mental involvement compared to laptop note-taking.

References

  1. The effects of writing on LSSM learning — The impact of various writing-to-learn activities.
  2. Repetition of Input and Notes — Effects of repetition on recall with note-taking.
  3. AI response to my question challenging the Luo study that assumed the notes themselves could not result in generative activity — How note modification strategies can counter the conclusions of Luo et al.’s study.

Output from RAG input of all notes

Hand vs. Keyboard Note-Taking: What Your Notes Show

Your notes contain extensive coverage of this topic across multiple studies.

Here is a synthesis of the key findings:

The Foundational Study

The landmark study by Mueller & Oppenheimer (2014), documented in Pen is mightier than the keyboard, found that keyboard users recorded more content but that longhand note-takers performed better, particularly on conceptual understanding items. In Study 3, which included a week delay with note review, differences favoring handwritten notes emerged specifically when review was allowed.

The Generative Processing Explanation

Several of your notes describe the primary theoretical explanation for a longhand advantage. Because typing speed (~33 wpm) exceeds handwriting speed (~22 wpm), as noted in transcription speeds, longhand note-takers are forced to summarize and paraphrase rather than transcribe verbatim. This engages generative processing-connecting new information to prior knowledge and reorganizing ideas-which promotes deeper learning. As described in Luo, Linlin, Kenneth A. Kiewra, Abraham E. Flanigan, and Markeya S. Peteranetz2018:

“Laptop note takers had a transcription orientation and recorded mainly verbal information… whereas longhand note takers had a generative orientation and used a combination of verbal and spatial note-taking strategies.”

The Luo et al. (2018) study further found that keyboard users recorded fewer visual notes (charts, graphs) and that the process function was more important for keyboard notes while the product function (review) was more important for longhand notes.

Contradictory and Nuanced Findings

Your notes reveal that the research is far from settled:

  • The 2022 Meta-Analysis (2022 Notetaking Meta-Analysis) by Voyer et al. analyzed 36 articles and found no overall effect for notetaking approach. The authors speculated that studies showing a longhand benefit may have been confounded by distraction opportunities on digital devices.
  • Urry et al. (2021) (Urry.et.al) replicated Mueller & Oppenheimer and found inconsistent results. Laptop users took more notes, and more notes were related to better test performance. Importantly, Urry differentiated quantity from verbatim overlap-quantity was positively related to performance while proportion of verbatim overlap was negatively correlated.
  • Note completeness research (NoteCompleteness) challenges the longhand advantage by showing that the volume and completeness of notes influence retention. Bui et al. (2013) found that students with access to more complete notes (as produced by keyboards) demonstrated better performance, directly contradicting the Mueller & Oppenheimer interpretation.

The Revision Factor

Your notes highlight that the comparison changes substantially when revision is introduced:

  • Flanigan et al. (2023) (Flaniganetal2023) found that among note revisers, longhand note-takers added three times as many complete ideas during revision pauses as computer note-takers, and scored more than half a letter grade higher on posttests. However, the advantage emerged primarily through the revision process rather than initial note-taking alone.
  • Cojean & Grand (2024) (Cojean&Grand2024) found that computer note-takers were more likely to reformat their notes after class, suggesting they compensate for shallow initial encoding through later revision-a finding that complicates simple hand-vs-keyboard comparisons.

Research with Children

Your note childrennotes describes Horbury & Edmonds (2021), which is notable as one of the few studies with children (ages 10–11). They found conceptual understanding was superior when taking notes by hand, but unlike adult studies, there were no significant differences in the amount of content recorded. The authors speculated that children’s lesser keyboarding experience may have produced greater cognitive load when typing.

Critical Perspectives from Your Notes

Your notes contain several important critiques of this research area:

  • The generative argument has a logical flaw: As you note in transcription speeds, the logic assumes “a strange lack of awareness on the part of the learner. It would seem that the keyboarding learner is oblivious to the benefits of generative processing and goes with what is easier, but less effective.”
  • Modification can equalize outcomes: Your AI response to my question challenging the Luo study that assumed the notes themselves could not result in generative activity argues persuasively that generative processing can be shifted to the review phase, that laptop notes provide a richer base for modification, and that laptop affordances enable superior reorganization. The conclusion: “The medium matters less than the strategy.”
  • Ecological validity concerns: Multiple notes flag that most studies use brief delays and immediate testing, which differs significantly from actual classroom conditions where review may occur days or weeks later (Luo, et al 2018, Flaniganetal2023).

Summary

FindingSupport
Longhand produces more generative/paraphrased notesConsistent across studies
Keyboard produces more complete/voluminous notesConsistent across studies
Longhand leads to better learning outcomesMixed—meta-analysis shows no overall effect
Note completeness predicts achievementSupported, which favors keyboard
Longhand advantage strongest with note reviewSupported in some studies, not replicated in others
Revision can compensate for shallow keyboard encodingSupported by modification research

The overall picture from your notes is that the hand-vs-keyboard question is far more nuanced than popular accounts suggest. The medium matters less than the strategy the learner employs, and the opportunity for revision and post-processing may be the most important factor in determining which approach is more effective.

Loading

Karpathy Plugin for Obsidian

I have spent a significant amount of time over the past week or so developing a Karpathy wiki based on a large portion of my Obsidian notes. This process began in April when I decided to purchase a Macintosh Mini I intended to devote exclusively to the exploration of AI on the desktop. I was a bit slow in making this purchase and it took until last week to receive my purchase. My tardiness cost several hundred dollars more than it would have a few months ago. 

I was motivated to invest and explore this area for two reasons. First, my main interest in AI continues to focus on retrieval-augemented generation (RAG) of the notes and highlights I have collected to serve as a foundation for my writing projects. As I have used AI plugins to interact with the content I have stored and organized in Obsidian, I discovered that the API-based services for interacting with these notes are relatively expensive because the process of first feeding the notes to the AI service must be repeated each time a session is initiated. Karpathy proposed that AI could be used to create a wiki based on the concepts and connections in a collection of source material, either once (when the collection was created) or as each new item of content was added, and this wiki could then be the focus of future explorations, reducing the cost due to the repeated input of the same content to the AI service. 

My second motivator was personal curiosity, sparked by the many posts promoting the potential of AI tools and models that could run on personal hardware, avoiding the costs and scrutiny associated with using online services from major AI companies. The proposal was that many common uses of AI no longer required access to $20 or $200-a-month subscription services. 

I understand that another tutorial or “how I did it” post may be required at this point, but I read a post explaining that getting started with self-hosting LLMs will not be easy as the posts newbies are likely to read make it sound, and a good deal of exploration and personalization will be required. The message was intended not to be discouraging but to communicate that “don’t give up, you should be able to get it to work.” This was pretty much my experience, and I thought it worthwhile to explain the issues I encountered and why I had to make adjustments to my specific situation. My experiences with tech since the mid 1980s have kind of gone this way. 

So, I have two Mac Minis now, and the first challenge was how to connect both to the same large monitor so I can switch back and forth as required by a general-use and a specific-use computer arrangement. I knew I would have to purchase a KVM (keyboard, video, mouse), but I had not considered that my current setup uses a Bluetooth mouse and keyboard. More specifically, Apple’s Magic Keyboard and Mouse are not intended to be linked to more than one device. You charge your Magic keyboard with a USB cable, so the cable can be used as it has long been used to connect to a computer. You also charge your Magic Mouse, but the cable is inserted on the bottom of the mouse, preventing it from being used while it is being charged. Solution – purchase a mouse with a cable. The first challenge is overcome.

My plan was to use the Obsidian Karpathy LLM wiki plugin because this seemed the most efficient way to create a working system. The plugin’s setup allows selecting multiple AI sources, including subscription services. I did use Anthropic’s Claude API when I was having difficulty getting either of the two local options (Ollama or LMStudio) to work. Claude worked great, but adding one new source document cost 70 cents. My present collection is close to 300 note files, and the work the AI does increases as the complexity of the wiki increases so I treated the success as a sign the struggles I was experiencing could eventually be overcome. 

When using Ollama, I was experiencing a consistent problem with some, but not all of the note files the AI was ingesting to build the wiki. I spent a considerable amount of time over several days comparing the files that could and could not be processed and I never did find a difference. It wasn’t the length, the presence of specific markdown tags, the tool I had used to create the original markdown file, or any other variable I could imagine. Nothing. However, the problem was consistent. The same files, time after time, would either work or fail. 

My typical strategy in such situations is to ask questions of the Internet. One proposal was that the JSON history had become corrupted. The solution was to reveal the invisible files (the .files and folders) and delete these files. New files would be generated when the Obsidian app was next launched. This was done without consequence.

One issue I encountered was that the models displayed as options within Ollama did not contain the model (qwen2.5) I had found recommended when I read the descriptions of others. I searched how to add other models to Ollama and found it could be done with a terminal command (ollama pull <model name>. Now qwen2.5 appeared. Qwen3.6 was originally listed and I assumed there would be little difference, but for some reason, I was wrong, and the system worked with qwen2.5. 

Without going into details because others have already provided tutorials, you first add and install the Karpathy LLM Wiki plugin for Obsidian. The gear icon associated with this community plugin provides a “fill in the blank” form where you enter information linking Obsidian to the AI online service or local option you want to use. 

The wiki construction process is controlled by the commands that appear in the Obsidian command list when the Karpathy plugin is installed.

So, you start Ollama and select the model you will use in Obsidian. Start Obsidian and select the command to ingest a file or folder and be patient. Eventually, your wiki will be generated, and you can query the wiki rather than the source files. The right-hand column displays a response to a prompt.

So, I was able to generate a wiki based on more than 150 of my notes. In examining some of the components of the wiki I did find some weird artifacts. There were some with Chinese characters. I happen to be listening to LeoLaporte talking about different AI models and he said that qwen originated in China. It then made sense to me that the model might translate some of the Chinese names in my article summaries and include their Chinese translations (no idea if that is actually what happened). I also found some md pages with titles, but no content.

When I used the Karpathy command to submit queries, I found the quality of the responses to vary. Some made sense and some ignored sources I knew existed and were central to what I expected. Rewording of the query in ways I thought the model should have understood as equivalent sometimes resulted in the response I expected. 

The following comparison (first from querying my new wiki, and the second from using Claude with my original source files) demonstrates fairly substantial differences. As you can tell from a quick skim, there are significant differences in the eloquence and completeness of the responses. The sample responses are presented as an Appendix.

Summary

Each user is likely to have a unique experience when applying a local AI system to their content to generate a Karpathy-style wiki. Expecting that it will take some tinkering is good advice even when following instructions or “how to do it” videos on YouTube. My effort here was to describe what this tinkering experience looked like. 

After at least getting the system to work, I was able to compare my first successful effort with what I generated for 44 cents using Claude. The differences were substantial.

I assume there are many more issues to explore. I am currently repeating the process with a different script that instructs the AI on what to do when creating the wiki. If that effort produces substantial improvements, I will have to write another post.

*******

Response of local model to wiki

Summary of Research Comparing Handwritten Notes vs. Computer Notes

According to the literature summarized from various studies, there are several key findings regarding the comparison between handwritten notes and notes taken on computers:

  1. Repetition in Note-Taking:

   – A study by Bui et al. 2015 suggests that repeated exposure to material enhances learning outcomes, especially when coupled with effective note-taking strategies.

  1. Handwritten Notes and Memory Retention:

   – Longhand notes are described as more paraphrased and containing more signals and images, which can lead to higher text-related achievement when reviewed (from Longhand Notes).

   – Lecture Note-Taking Strategies indicate that students who take notes by hand outperform those who use digital devices, as handwritten notes require a higher level of processing and engagement.

  1. Editing and Revisions:

   – Research by Mueller and Oppenheimer (2014) found that while typing is faster, it leads to less cognitive effort and thus less retention due to the lack of engagement required in handwriting notes.

   – Follow-up revisions are crucial as they enhance memory retention. According to a study cited by AI response, revising digital notes can mitigate some of the disadvantages of digital note-taking, improving long-term recall and academic performance (from AI Response to Challenging Luo Study).

  1. Professional Note-Taking:

   – Additionally, in the context of professional training, effective note-taking is vital for future practitioners (from “Notetaking in Professions”).

References

  1. Longhand Notes — Core mechanism explaining why handwritten notes are more beneficial.
  2. Bui et al. 2015 — Effects of Repetition on Note-Taking and Recall Strategies in College Lectures.
  3. “Notetaking in Professions”. Stacy and Cain, American Journal of Pharmaceutical Education (2015) — Importance of note-taking skills for professional training.
  4. Mueller and Oppenheimer (2014) — Pen and Pad versus Laptop for Note-Taking.
  5. AI Response to Challenging Luo Study — How Note Modification Research Might Counter Luo et al. (2018)

*******

Paste from Claude plugin within Obsidian

Handwritten vs. Computer Note-Taking: What Your Notes Say

Your vault contains extensive research on this topic, and the picture that emerges is more nuanced and contested than the popular narrative suggests. Here is a synthesis of what your notes reveal:

The Foundational Study

The most influential study is Mueller & Oppenheimer (2014), documented in Pen is mightier than the keyboard. Their key finding was that faster keyboard transcription led to more verbatim, less generative notes, while the slower speed of handwriting forced summarization and paraphrasing. Notably, their Study 3 found no differences when note review was not allowed, but longhand was superior when review was allowed—a surprising result.

The Generative Processing Argument

Several of your notes outline the theoretical rationale:

  • Translation hypothesis: Writing adds a distinctive encoding channel beyond just listening (Luo, Linlin, Kenneth A. Kiewra, Abraham E. Flanigan, and Markeya S. Peteranetz2018)
  • Generative hypothesis: Handwriting’s slower speed forces paraphrasing, organization, and integration with prior knowledge, leading to deeper processing (Depth and Generative, EncodingEffect)
  • As noted in transcription speeds, adults type at ~33 wpm and write longhand at ~22 wpm, meaning handwriters must be more selective

Your note in transcription speeds raises a sharp critique of this logic: it “suggests a strange lack of awareness on the part of the learner. It would seem that the keyboarding learner is oblivious to the benefits of generative processing and goes with what is easier, but less effective.”

Contradictory and Inconsistent Findings

Your notes consistently emphasize that research findings in this area are contradictory:

  • Urry.et.al documents a major replication of Mueller & Oppenheimer that found inconsistent results. Urry found laptop users took more notes and more notes related to better performance. Critically, they differentiated quantity from verbatim overlap—quantity was positively related to performance while verbatim proportion was negatively correlated.
  • Luo, et al 2018 found that when notes were recorded but not reviewed (process function), laptop was actually better for image-related learning. Only when notes were reviewed (product function) did longhand show an advantage. The authors themselves note that “findings in this area of research are often contradictory and uncontrolled variables may be responsible.”
  • Flaniganetal2023 found that, surprisingly, longhand and computer note takers recorded a similar number of ideas during the lesson, contradicting the expected computer advantage. However, longhand revisers added three times as many complete ideas during revision pauses.

The Note Completeness Counter-Argument

Your notes in NoteCompleteness and Colllaborative Notes document an important counter-narrative: note completeness predicts achievement, and computers produce more complete notes. As noted:

“The studies of note volume and completeness would seem to challenge the argument that taking notes on a computer is a liability because more notes are taken implying less thinking about the content.”

Bui et al. (2013) found that students with access to their more complete (computer-generated) notes demonstrated better performance, directly contradicting Mueller & Oppenheimer’s interpretation.

The Revision Factor

Multiple notes highlight that the handwriting vs. computer debate may miss the most important stage entirely—revision:

  • Three stage model of note-taking proposes that effective note-taking is a three-stage process: recording ? revision ? review, and most studies ignore the revision stage.
  • Cojeanatal2024 found that computer note-takers were more likely to reformat and revise their notes after class, potentially compensating for less generative initial encoding.
  • Flaniganetal2023 found longhand revisers added more ideas during revision pauses, but this study used interspersed pauses—an artificial design unlikely to reflect real practice.
  • Cohen2013ActiveSummary demonstrated that note restructuring significantly improved exam scores regardless of medium, suggesting what you do after taking notes matters more than how you take them.

The Children’s Research

Your note childrennotes documents a unique study (Horbury & Edmonds, 2021) with 10–11-year-olds that found:

  • Conceptual understanding was superior with handwritten notes
  • Unlike adult studies, there were no significant differences in note quantity between conditions
  • The authors speculated that children’s limited keyboarding experience may have increased cognitive load when typing, a different mechanism than proposed for adults

Your AI-Generated Synthesis

Your note AI response to my question challenging the Luo study that assumed the notes themselves could not result in generative activity offers a compelling reframing:

Luo et al. ConclusionModification Research Counter
Laptop notes are transcription-orientedGenerative processing can be added during revision
Longhand notes are better review productsLaptop notes can be transformed into superior products through revision
Laptop notes lack images and signalsThese can be added easily using laptop affordances
Reviewing longhand notes yields higher achievementStudies only tested passive review, not active modification
Longhand is generally preferableThe medium matters less than the strategy

Overall Conclusions From Your Notes

  • The popular narrative favoring handwriting is overstated. Replication attempts and meta-analyses show inconsistent results.
  • The mechanism matters: Quantity of notes is positively related to achievement; verbatim overlap is negatively related. These are separate variables that should not be conflated.
  • Most studies use artificial conditions (immediate testing, brief delays, no real revision opportunities) that don’t reflect actual academic practice.
  • Revision and post-processing may be more important than the initial recording medium, and computers offer superior affordances for revision.
  • The real issue is strategy, not medium. Untrained students default to transcription on keyboards and forced summarization by hand—but training could equalize outcomes.
  • Note completeness matters for the storage function, and the storage function is generally more important than the encoding function (ExternalStorage).

Loading

Digital Devices and Effective Studying and Long-Term Note Use

Lately, I seem to encounter a great deal of advice warning against reading and note-taking on a device. Paper and a pencil are argued to be better. In a recent post, I argued that I found these complaints ill-advised. This is a second attempt to present the same argument and a different way to justify my position.

Most simply put, I don’t think the researchers and I are talking about the same thing, even though it seems the anti-tech types would make the case that we are. The research frequently cited reading and taking notes tend to focus on more focused and short-term processes than what I have in mind. My focus, and I think the actual focus of most learners, is more on what I might describe as studying or the use of notes as imagined by the PKM/Second Brain aficionados. When I read a fiction book for pleasure, I might have a better understanding of that book if I read the content on paper. This is very different from how I would best get from reading a textbook to doing well on an exam a couple of months from now. A similar comparison might be made with taking notes during a lecture. I would seldom be taking notes for an exam the same day, but more likely for one week away. 

Process Models

I write a lot about process models relevant to understanding and developing both learning skills and the knowledge that results from applying those skills. The process model of writing (Flower & Hayes) makes a good example. These researchers proposed that a process model of writing was useful to researchers because the model identified subskills that could be studied to see how these contributing skills might explain the performance of more and less effective writers and to educators trying to understand what contributing behaviors might be isolated for practice and development. 

I suggest that the same type of process model would be helpful for developing study skills and for taking a different look at the possible advantages of using digital devices for reading and note-taking. The argument in this second case is that research comparing tech vs. traditional approaches has overlooked important processes in studying and note-taking applications.

Processes in tasks involving the collection and eventual application of information

There have been efforts to identify the processes involved in translating a presentation (e.g., a lecture or book chapter) into intended applications. Several researchers (e.g., Cojean & Grand, 2024; Flanigan et al., 2023; Luo et al., 2016) have extended the original two-stage model (note-taking and external storage) to emphasize the importance of revision. Returning to the importance of multi-process models in understanding the potential issue of whether it matters if one takes notes on paper or using a device, the studies differ. Flanigan and colleagues engaged the unusual practice of inserting pauses during a presentation to allow for revision and found that those taking notes by hand created more revisions. In contrast, Cojean and Grand found that after class those taking notes on a device made more revisions. Systems of taking notes, for example Cornell Notes (Pauk & Owens, 2011) and recent PKM systems (Ahrens, 2022; Forte, 2022), differentiate revision as a separate process in the use of notes. 

In the spirit of the writing process model, I have created my own identification of note-taking and note-using processes listed as a sequence with the recognition that notetakers frequently revisit earlier processes after finding a limitation in what a later process makes available. The sequence of descriptors for these processes follows. 

  • Collecting
  • Considering
  • Elaborating 
  • Exporting

Collecting – creating a representation of content (presentations, videos, text material) for use in the future

Examples – creating annotations, notes, highlights

Considering  – offline processing of the information collected for personal understanding and to evaluate gaps in understanding

Examples – rewrite existing notes based on comparison of personal collection with that of peers, return to source material to fill in gaps

Elaborating – speculation based on personal understanding of original information for fit within existing knowledge and potential application

Examples – links to existing notes on similar topics, Internet searches to locate and augment existing notes with additional examples of key concepts

Exporting – use of cumulative stored content to meet personal or assigned goals

      Examples – test performance, assigned writing tasks, personal writing projects

The Processes and The Question of Handwriting vs. Digital

My contention is that when tasks involve all of the processes I have identified, digital tools offer advantages in the efficiency of collection, storage, search, and manipulation. These advantages are magnified when the task’s time frame is extended and initial goals are unclear. I have written at length on these topics and I have tried to organize some of these posts, organized by process, below. I have avoided considering how AI might be used in these processes, but such engagement would be far easier if working in a digital environment. 

Collecting

Take digital notes for best lecture performance

Note and highlight extraction for efficient review and storage (Readwise for books, Highlights for PDFs)

Considering

Note and highlight extraction for efficient review and storage (Readwise for books, Highlights for PDFs)

The Power of Collaboration: Enhancing Your Note-Taking Experience

Preserving context in digital writing

Elaborating

Smart Connections finds note connections

Highlighting in the age of digital content

Notes and the Translation Process

The Space Between Encountering Information and Application

Digital for serious reading tasks

School and Professional Note-Taking

Exporting

School and Professional Note-Taking

Resources

Ahrens, S. (2022). How to take smart notes: One simple technique to boost writing, learning and thinking

Cojean, S., &  Grand, M. (2024). Note-taking by university students on paper or a computer: Strategies during initial note-taking and revision. British Journal of Educational Psychology, 94, 557–570. https://doi.org/10.1111/bjep.12663

Flanigan, A. E., Kiewra, K. A., Lu, J., & Dzhuraev, D. (2023). Computer versus longhand note-taking: Influence of revision. Instructional Science, 51(2), 251-284

Forte, T. (2022). Building a second brain: A proven method to organize your digital life and unlock your creative potential. Simon and Schuster.

Luo, L., Kiewra, K. A., & Samuelson, L. (2016). Revising lecture notes: How revision, pauses, and partners affect note-taking and achievement. Instructional Science, 44(1), 45-67.

Pauk, W., & Owens, R, (2011). How to study in college. Boston, MA: Wadsworth, Cengage Learning.

Loading

Cooperative Learning When AI Is Your Partner

Various terms have been used to describe AI and human partnerships in attempting to accomplish a goal defined by the human. For example, a recent book by Ethan Mollick was titled “Co-intelligence”.  Recently, I have been reading about the work of several authors who have described different ways in which learners might interact with AI some being more successful than others. I will return to this material after my preliminary remarks. These analyses also consider various types of collaboration.

As I read the most recent set of papers, I flashed back on research I encountered in the early 1990s. In this case, there was no AI partner, but educational researchers and instructional designers were investigating ways in which student peers could collaborate (e.g., Johnson, et al., 1991; Slavin, 1995). If you were a preservice or practicing teacher at that time, you likely heard a lot about cooperative learning. There were multiple proposed benefits of cooperative learning. Social interactions were motivating. Multiple individuals have unique experiences and skills and combining these resources benefits all who are exposed. Interaction, whether it be a form of teaching, working through differences of opinion or error-checking each other, or simply sharing experiences, augments an individual’s cognitive activity. More individuals, theoretically, allow more to be accomplished in less time. 

There were also concerns about cooperative learning activities, and if the connection to learning with an AI partner is not obvious, identifying educators’ and researchers’ concerns should clarify the similarities. The first was called the freeloader effect – in a pair version of cooperative learning, one participant might do all of the work and the other would do and learn very little. Perhaps one student would simply be more motivated or more capable and find that working alone was more efficient. Other concerns included a lack of experience and skill in cooperative planning or in effectively using the talents of multiple individuals. A few remedies from that era I recall included individual accountability (e.g., individual tests on content), positive interdependence (e.g., clear identification of tasks whose outputs will be combined in the final product), and a structured or scaffolded process. 

Scaffolding will come up again when discussing the potentially effective use of AI, so perhaps an example of what this might look like would be helpful. In building construction, a scaffold provides temporary support for workers. In education, a scaffold provides a structure that supports a learning task by guiding how work is done. Consider a version of a common strategy for a cooperative task – i.e., think, pair, share. Students are given a task. Then, each student writes (or just thinks) of a proposed solution. Finally, students consider each other’s proposals (pair) and integrate them to arrive at a solution (share). The imposed structure in this case ensures that each individual participates and provides a record of their work, should the teacher want to hold each accountable for participating. The jigsaw cooperative learning technique offers another method for scaffolding a cooperative learning activity to promote participation and individual accountability. A project is selected that requires identifiable roles or tasks. For example, when creating a brochure describing the butterflies one might most likely encounter in a local garden, individual students could be assigned to research different butterflies and then asked to combine their research into a single document. 

Before moving on to the cooperation between a student and AI, I propose that a body of research examines peer interaction in educational settings, and that its identified issues and remedies might be useful to those now focused on AI. 

Differentiating Ways Learners Use AI In Attempting to Identify Productive Approaches

First, I would refer readers to a previous post in which I examined the classification scheme of AI learning strategies. The following is my previous explanation of the levels in this classification scheme. 

The three levels have the following characteristics:

Zone 1: No AI Involvement

In this level, learning occurs without any AI assistance. While learning happens, it is often “capacity-constrained” because the learner must spend significant time and effort on execution and task completion, leaving less bandwidth for higher-order reflection.

Zone 2: Scattered, Half-Hearted Use

    This is characterized by using AI for minor tasks like fixing sentences, checking facts, or tidying paragraphs. It often produces the worst learning outcomes. The learner still carries nearly the full cognitive load but adds the overhead of managing AI interactions without gaining significant cognitive savings. Note: this summary paraphrases the description of the authors. My version would add having the student using the AI tool to perform the task based on simplistic instructions. 

Zone 3: Committed, Strategic Delegation

    This level involves offloading entire categories of substantive work to AI to free up genuine cognitive capacity. This freed bandwidth is then redirected toward tasks AI cannot do, such as critiquing frameworks, questioning assumptions, and making complex judgment calls. This zone is where “transformative learning” is thought to live, provided the course design is intentional about how and why tasks are delegated.

My attempt to summarize this scheme suggested that the use or nonuse of AI and its relationship to successful learning experiences could be explained by investigating how certain conditions of student motivation, metacognitive proficiency, and working memory interact. For motivation, was the learner focused on developing personal knowledge and/or skills beyond task completion and on receiving a positive evaluation? Working memory reflects the capacity for meaningful learning beyond underlying task demands. The notion here is that AI might, in some situations, handle nonessential or untargeted tasks, allowing the learner to devote their attention and processing capacity to accomplishing targeted goals. Metacognitive proficiency suggests that more sophisticated learners with sufficient available attentional capacity are more likely to make sound decisions about when to use AI to free cognitive capacity to accomplish goal-related knowledge and skills. 

I hope it makes some sense how these factors might interact. Allow an argument based on a personal perspective. I would suggest that my own learning offers motivational advantages. I learn to accomplish personal goals rather than be subject to external goals and reward structures in a classroom setting. I am also more metacognitively sophisticated than secondary school students. I understand the tasks I want to accomplish well having explored them countless times over the decades and such experiences offer me useful insights, but also means that I have background knowledge and cognitive skills that require less working memory capacity on my part. My use of AI could enable Zone 3 processing. I am not saying that this is always the case, but it makes sense I have a greater opportunity to function at this level.

Does this analysis then suggest that less-experienced learners, even secondary students, should not work with an AI partner? Not necessarily. This is where the scaffolding found to be important in cooperative learning and proposed as a way for Zone 3 functioning to be practical. Scaffolding bridges the gap, allowing tasks to be accomplished before all necessary conditions are in place and offering a mechanism for introducing required skills.

The mention of design in the Zone Three descriptions is another way of suggesting the value of scaffolding. I have included multiple references I found that explain the three-zone model and offer suggestions for scaffolding. My typical reaction is that the examples never seem to include the tasks educators might most need assistance in addressing. For example, writing tasks are commonly described, but not general “homework” tasks. 

I think the best advice is to focus on processes and discuss assignment goals with students, differentiating those processes students are free to use AI to accomplish and which are expected to be completed by students. Consider how students might document their activity in completing each. For example, again using a writing example, use AI to identify content you will use in your writing project (available for submission). Submit your draft based on this content. Have AI evaluate the quality of your draft (available for submission). Submit your final version. Students might also be given a general topic on which a written product will be required. Students will bring the relevant content they have found to class and then receive specific instructions on the product to be fashioned during that class period. Both resources and product to be submitted. 

General Resources:

Mollick, E. (2024). Co-intelligence: Living and working with AI. Penguin. Penguin.

Johnson, D., Johnson, R., & Holubec, E. (1991). Cooperation in the classroom (rev. ed.). Edina, MN: Interaction.

Slavin, R. (1995). Cooperative learning: Theory, research, and practice (2nd ed.). New York: Merrill.

Sources for AI Level Analysis

Hardman – The cognitive offloading paradox

Lodge and Lobel. Artificial intelligence, cognitive offloading and implications for education

Means. Strategic Cognitive Offloading: What the Research Says, and Why Higher Education Isn’t Ready for It

Wang, S., Zhang, H. Pedagogical partnerships with generative AI in higher education: how dual cognitive pathways paradoxically enable transformative learning. Int J Educ Technol High Educ 23, 11 (2026). https://doi.org/10.1186/s41239-026-00585-x

Loading