AI (Grammarly) and Writing: Good or Evil?

I am interested in the potential of AI in developing writing skills. I am presently focused on Grammarly, having used the tool for years, and am now considering how it might play a productive role in secondary and higher education efforts to develop writing skills. Exploring what those writing about Grammarly on Medium have to say about Grammarly, I have come away with the impression the majority argue the tool is overpriced, less effective than tools with a similar purpose, and generally a bad idea when applied in an educational setting. I don’t agree. 

Writing in the classroom

I like to draw a distinction between learning to write and writing to learn. This distinction is artificial, as classroom instructional strategies such as “Writing Across the Curriculum” argue that both goals can be addressed when writing assignments in other disciplines are evaluated both for the quality of the writing and for what the writing suggests about students’ understanding of a given topic. My argument here focuses on the potential value of AI in learning to write. 

When and why is AI a problem when learning to write

In situations where the development of writing skills is the emphasis, an AI tool is argued to be problematic because students cheat by using AI to avoid doing work that requires them to practice the skills they are expected to master. In addition, by turning in work for evaluation that students did not actually perform, they do not receive feedback on the skills they are supposed to be learning and are credited with achievements they have not actually earned. 

AI and writing: A different take on the actual problem?

Many educators, aware of the possibility of cheating, have resorted to approaches such as short, handwritten in-class assignments that eliminate the possibility of using AI. There are limitations to this approach, especially for the unique skills required to create longer arguments or other lengthier projects. 

As adults with reasons to write and without worrying about the need to prove everything that appears in a written product is based on our own knowledge and writing skills, we may take advantage of AI in many different ways. One real question is how, and perhaps if, we are preparing students to transition from a focus on learning new skills or the graded demonstration of one’s knowledge to a combination of AI and personal knowledge and writing skills. In addition, some are suggesting, and again rightfully so, that AI can benefit students’ efforts to learn to write and write to learn. Here, I want to emphasize the potential assistance in learning to write more effectively. 

I tend to react to what I think are naive expectations of teachers and the reality of working in classrooms is important here. For example, I support the exploration of AI as a tutor, not because I think AI is equivalent to a human tutor, but because human tutoring is costly and many students who need help do not receive sufficient human attention as a consequence. I have a similar opinion about learning to write. It would be great if each student could write a lot and receive rapid feedback, as well as an individual conference related to their effort. Neither immediate, consistent feedback nor frequent individual attention is practical. Just having an AI writing tool, such as Grammarly, that can provide immediate feedback on what has been written seems like a practical improvement.

So Much Depends on Personal Motivation

Grammarly and asking pretty much any AI tool to evaluate specific attributes of your writing quality will provide you with feedback to consider. The issue is really whether you take the time to ask for this feedback and to consider the feedback that is produced. Here is what I mean. I use Grammarly while I write, and it constantly provides feedback. In reflecting on my own behavior, I almost always quickly accept the suggestions for what I have written (these appear as underlines in various colors) by clicking to have Grammarly fix the problem. I don’t stop to figure out what was wrong with what I wrote. Was that an actual error of grammar or spelling, and if so, why? The fixes always seem better, but they also remove what may just be my voice or personal preference in how I say something. I avoid the opportunity to learn and also allow Grammarly to “standardize” my writing. At this moment, admitting this has made me self-conscious. 

This reminds me of the experience I had providing comments on many of my grad students’ theses and dissertations. In later years, I liked to use the comments feature in Google Docs to leave comments and identify actual errors. I started to realize that some students were simply allowing me to rewrite their papers, when what I wanted was for them to consider something different. Often, I had to remind them of the difference between my thoughts about their work and the actual errors I pointed out. 

If you use a tool such as Grammarly, you probably recognize my observation in your own behavior. It is so easy to accept proposed changes based on a kind of “that sounds pretty good thinking” and trust in the assumption that the “system knows the rules better than I.” Taking this approach is quick, painless, and “good enough.” The problem is that this approach fails to take advantage of at least some of these situations to learn. Why were these changes recommended? Is my way of expressing myself flawed or just unique? Grammarly will help you consider which is most likely. 

What was wrong with what I wrote?

Grammarly has always allowed you to pause when suggesting a change. There was no time limit on the opportunity to consider what you wrote in comparison to what was recommended. As the tool was improved and with the more recent integration of AI, efforts were made to explain why a change was recommended. At first, the tool offered a general reference to rules. Here is what a split infinitive is, and here are some examples of sentences containing a split infinitive and improved versions of the same sentences. Here is an example of passive voice, and here are some examples. The most recent advance offers similar information, but specifically related to your own words rather than just generic examples. 

One note – I have encountered descriptions of this newest capability from others, but I haven’t been able to replicate the same output on my own computer with the latest version of Grammarly. I have had this difficulty even though I input exactly the same text used in the other demonstrations I have encountered. My setup will identify the error and provide generic examples, but it won’t explain based on the text I have entered. I can generate explanations specific to my written text, but I have to use the AI window to enter a prompt asking for this information (see examples below). 

Here are a couple of examples. In the first, you see a sentence with three components underlined in blue (I highlighted it in blue so you can find it). In the associated column on the right you see the proposed alternative with the changed words or punctuation bolded. The red box identifies the button to get additional information. The second image shows the result of making use of this button. The explanation for the proposed changes – Our revised version of your text may make it more effective. – is Grammarly’s explanation. As I said, this offers little of value to me. However, if I use the AI prompt box (Grammarly calls this Superhuman Go), which appears at the base of the right-hand column, I can ask for an explanation, and I receive the lengthy explanation shown in the third image. 

Here is a second example that explains a split infinitive. Again, if I ask for an explanation related to the limitation in my writing, I get both a general and a targeted explanation of the split infinitive (second image).

Improvements at multiple levels

Improvement in writing happens at multiple levels. There are spelling errors and grammatical failures. These are the easiest for a digital tool to identify, and spelling correction has been around for decades. These might be described as superficial errors. There are deeper levels of good writing that any of us who critique others’ work find much harder to evaluate and that are the most difficult proficiencies to teach and learn. Grammarly’s AI capabilities provide a way to address these proficiencies. The easiest way I can describe this capability is to list some of the prompts Grammarly suggests. 

  • “Critique the logic and reasoning in this text.” 
  • “Are there any gaps in my argument here?” 
  • “Does this conclusion follow from my premises?” 
  • “Point out any weak points in my reasoning.” 

I will append Grammarly’s extensive analysis of a recent post based on one of these prompts (see Appendix). I have extensive experience as an academic writer with two published textbooks. These works involved feedback from a dedicated editor and input from multiple paid professors regarding both the content and communication style. The detail in Grammarly’s feedback is easily as detailed. 

Summary

The point I am trying to make here is not that AI cannot be used to cheat. My point is that Grammarly and other AI tools can also provide feedback more quickly and consistently than is practical for educators. This seems to be another example of the constant dilemma with AI, it can be used for good or for evil. In this case, when a writer is working on their own, they decide which it will be. 

Addendum

Grammarly’s analysis of AI Processing Options

Prompt: Point out any weak points in my reasoning

A few places in the reasoning could be stronger or more precise:

  1. The continuum needs clearer criteria
    You describe a continuum from “total AI” to “total human,” but the dimensions of that continuum are a little mixed. Sometimes the distinction is about who selects the source material, sometimes who reads it, sometimes who organizes it, and sometimes who writes the final text. You might strengthen the piece by naming the dimensions explicitly, such as:
  • control over source selection
  • depth of human engagement with sources
  • human contribution to interpretation
  • AI role in organization
  • AI role in final wording
  1. “Information input” may be too narrow
    In the Recall.ai example, you say the main thing you control is the information input. But choosing 15 PDFs is not just supplying information; it also frames the argument by deciding what evidence counts. That choice may involve substantial judgment, even if you do not read the documents closely.
  2. The quality question needs more defined standards
    You ask which approach produces the highest-quality product, but “quality” could mean many things: factual accuracy, originality, usefulness, writing style, source integration, depth of understanding, efficiency, or ethical acceptability. The argument would be stronger if you noted that different approaches might score differently depending on the quality standard.
  3. The “total AI” end may not be truly total AI
    Even in the ChatGPT/Claude example, the human still defines the topic, purpose, and prompt. So it may be more accurate to call this “minimal human input” rather than “total AI.”
  4. The “total human” end still includes tool-mediated thinking
    The Obsidian-only example is described as having no AI use, which makes sense if the point is AI involvement. But it still involves a digital tool that shapes organization, retrieval, linking, and thinking. If your continuum is specifically about AI involvement, that is fine—but you may want to clarify that it is not a continuum from tool-free writing to machine-written writing.
  5. The RAG example may understate verification
    You note that you do not have to read the documents, but that raises an important issue: without reading them, you may not be able to evaluate whether the AI’s summaries are accurate or whether the final product fairly represents the sources. This is an important weakness in that approach and could be made more central.
  6. The Karpathy Wiki stage could use more explanation
    This is the most complex example, but the reasoning around why it is distinct could be clearer. Is the key difference that AI creates a reusable knowledge structure? That it separates source processing from final generation? That it mimics a human note system? Stating the main conceptual difference directly would help.
  7. “Parallel to how human writers function” is promising but underdeveloped
    You suggest that identifying themes and concepts before drafting may parallel human writing. That is an interesting claim, but it needs a bit more support. Human writers do not just extract concepts; they evaluate, synthesize, prioritize, and sometimes change direction. You could strengthen this by explaining exactly which parts of the AI process resemble human writing and which do not.
  8. The ethical/authorship issue is implied but not addressed
    Your categories raise questions about authorship, responsibility, and acceptable AI assistance, but those issues are not directly discussed. If this is meant to help readers evaluate AI writing practices, you may want to include a brief note about where responsibility for claims and interpretations resides.
  9. The examples could be labeled more consistently
    Some category labels name tools, while others name processes. For example, “ChatGPT or Claude” names tools, while “Obsidian to Karpathy Wiki to Product” names a workflow. You might make the structure more analytical by using labels like:
  • Prompt-only generation
  • Source-constrained generation
  • Human-annotated source generation
  • AI-organized knowledge-base generation
  • Human-authored synthesis

Overall, the central idea is strong: you are not treating “AI writing” as one thing, but as a range of workflows with different levels of human involvement. The main improvement would be to make the continuum’s criteria explicit so readers can see exactly what changes from one category to the next.

Loading