Files
last30days-skill/docs/comparison-results/evaluation/eval-3-macbook.md
T
Matt Van Horn bed0557b65 feat(quality): GOAT synthesis improvements - hybrid cross-source linking, YouTube synonyms, human-readable xref tags
Ran 15-way blinded comparison (5 topics x 3 versions). CROSS won all 5 topics
(4.74/5.0 avg vs HN 4.10, Base 3.73). Then improved CROSS further:

- dedupe.py: hybrid similarity (token+trigram Jaccard) at 0.40 threshold,
  cross-source links went from 3 to 13 items across 5 topics
- render.py: [xref: HN5, HN4] -> [also on: HN, Reddit] for human-readable tags
- youtube_yt.py: SYNONYMS dict so "hip hop" matches "rap" (0.33 -> 0.71 score)
- SKILL.md: instruction #7 tells Claude to lead with cross-platform signals

Validation: improved CROSS scores 4.38/5.0 vs original 3.98 (+0.40), wins 4/5
topics. Biggest gains in specificity (+0.8) and format compliance (+1.0).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-25 16:06:53 -08:00

15 KiB

Evaluation: M4 MacBook Pro review

Query Type: RECOMMENDATIONS Label Map (REVEAL AFTER SCORING): {'A': 'hn', 'B': 'base', 'C': 'cross'}

Evaluation Rubric

Score each version 1-5 on these dimensions:

1. GROUNDEDNESS (30%)

Does the narrative cite specific sources from the research data?

  • 1: Generic statements, no citations, could be written without any research
  • 3: Some citations but mixed with pre-existing knowledge filler
  • 5: Every finding backed by a specific source ("per @handle", "per r/sub", "per [channel]")

2. SPECIFICITY (25%)

Are findings specific (named entities, exact numbers) or vague?

  • 1: Vague generalities ("AI video tools are improving", "developers are debating frameworks")
  • 3: Some specifics mixed with generic padding
  • 5: Named products, exact numbers, version names ("Seedance 2.0 added lip sync", "698 likes")

3. COVERAGE (20%)

Does the synthesis represent findings from all available data sources?

  • 1: Only mentions 1-2 sources, ignores others
  • 3: Mentions most sources but unevenly weighted
  • 5: Naturally weaves Reddit, X, YouTube (and HN if available) into the narrative

4. ACTIONABILITY (15%)

Does the invitation give specific, research-derived next steps?

  • 1: Generic "let me know if you want more info"
  • 3: Somewhat specific but not clearly grounded in research findings
  • 5: Each suggestion references a specific thing from the research ("I can compare Seedance 2.0 vs Kling")

5. FORMAT COMPLIANCE (10%)

Does it follow the expected output format?

  • 1: Missing stats block, no invitation, wrong structure
  • 3: Partial stats block, generic invitation
  • 5: Perfect stats block with real counts, source box-drawing chars, top voices identified

VERSION A

What I learned

The M4 MacBook Pro has been extensively reviewed across YouTube and Reddit, but the developer-focused Hacker News community has been notably silent on the topic this cycle. Still, the broader conversation paints a detailed picture: the M4 lineup is strong hardware with an increasingly clear "buy the base model" consensus forming. Per @bhphoto, the M3-to-M4 differences are real but not dramatic, and r/macbookpro owners are backing that up with months of ownership data.

Base M4 is the default recommendation. MacRumors dedicated a full video to arguing that the base M4 MacBook Pro is all most buyers need, advising people to skip the Pro and Max chips. Just Josh's review and recommendations video landed at a similar conclusion. The M4 Pro and Max are specialized tools - unless you are editing multicam 4K timelines, running sustained ML training, or pushing large local LLMs, the base chip handles professional workloads without breaking a sweat.

Battery reality is a sore spot. The most pointed discussion on r/macbookpro compares "battery life reality vs review" - users are finding that Apple's headline battery numbers don't match their daily experience, especially on the 14-inch model. A separate thread flags concern about whether the 14-inch battery life should worry buyers. The gap between controlled reviewer testing and real usage with Chrome, Slack, and Docker running is meaningful enough to generate repeat threads.

Long-term ownership tells the real story. The launch-day review cycle was overwhelmingly positive, but the 3-to-12-month follow-ups are more measured. Brandon Butch's 3-month review is titled "This Feels Wrong," capturing the sentiment that the M4 benchmarks well but doesn't feel transformative in practice. Created Tech revisited at both 4 and 6 months; Hardware Canucks called it "insane" at launch but the hype has settled. Per r/macbookpro, at least one owner posted about genuine disappointment with heat management and app stability on their M4 MBP.

Local LLM capability is a real differentiator. On X, the conversation has shifted toward the M4 as a local AI machine. @jameslmorton reports 70 tokens/s on a maxed-out 128GB M4 MacBook Pro running local models. @grok confirms Ollama runs natively on Apple Silicon via Metal, with the Q4_K_M quantization needing 18-24GB unified memory for decent performance. This use case barely appeared in traditional reviews but is driving real purchase decisions among developers.

The upgrade calculus from M3 is thin. @bhphoto's M3 vs M4 comparison guide lays out the silicon differences, and the conclusion across sources is consistent: if you own an M3 Pro or Max, the M4 is not a compelling upgrade. New buyers benefit; upgraders should probably wait for M5.

---All agents reported back! ├─ 🟠 Reddit: 3 threads │ 0 upvotes │ 0 comments ├─ 🔵 X: 12 posts │ 10 likes │ 1 repost ├─ 🔴 YouTube: 10 videos │ 7,866,700 views │ 3 with transcripts ├─ 🟡 HN: 0 stories (no results this cycle) ├─ 🌐 Web: 0 pages (supplementary) └─ 🗣️ Top voices: @bhphoto (3 likes), @jameslmorton │ r/macbookpro

---I'm now an expert on M4 MacBook Pro reviews. Some things I can help with:

  • Help you decide between the base M4, M4 Pro, and M4 Max based on your actual workflow
  • Estimate realistic battery life for your usage pattern vs Apple's marketing claims
  • Assess whether upgrading from an M1/M2/M3 MacBook Pro is worth it right now

VERSION B

What I learned

The M4 MacBook Pro has been out long enough that the conversation has shifted from launch hype to real-world ownership reports - and the picture is more nuanced than the initial review wave suggested. Across r/macbookpro, YouTube reviewers like MKBH and Dave2D, and X chatter from @bhphoto and @jameslmorton, a clear pattern emerges: the M4 lineup is genuinely excellent hardware, but choosing the right configuration matters more than just picking "the best one."

The base M4 is the real story. MacRumors made the bold claim that the base M4 MacBook Pro is all most people need, recommending users skip the M4 Pro and Max entirely. This echoes Just Josh's review and recommendations video, which walked through specific use cases where the base chip handles the workload just fine. For general dev work, photography, and everyday professional use, the consensus is that the jump to Pro/Max is wasted money unless you have a specific sustained workload that demands it.

Battery life claims don't hold up for everyone. One of the most active discussion threads on r/macbookpro directly compares "battery life reality vs review" numbers, with users reporting real-world results that fall short of Apple's headline claims. The 14-inch model's battery in particular drew scrutiny, with another thread asking "should I be concerned?" after seeing lower-than-expected drain rates. Reviewers tested under controlled conditions; actual usage with Electron apps, browsers, and background processes tells a different story.

Long-term reviews are more honest than launch reviews. Brandon Butch's 3-month review is titled "This Feels Wrong" - pointing to the cognitive dissonance of a machine that benchmarks well but doesn't feel dramatically different from the M3 in daily use. Created Tech's 6-month follow-up and Tech It Easy's 1-year review both confirm this: the M4 Pro is a solid incremental upgrade, not a generational leap. Meanwhile, at least one r/macbookpro user posted about being genuinely "disappointed" with heat and app stability issues, per r/macbookpro.

Local AI workloads are the new benchmark. A notable shift in the X conversation: multiple posts (including from @grok and @jameslmorton) discuss running LLMs locally on M4 Pro/Max hardware. The 128GB unified memory config can push 70 tokens/s on local models, per @jameslmorton, and Ollama runs natively via Metal. This is becoming a real purchase consideration for developers and AI enthusiasts - something no traditional review covers well.

M3 vs M4 upgrade: marginal for most. @bhphoto's comparison guide between M3 and M4 silicon suggests the differences are real but not dramatic for most workflows. Unless you specifically need the Nano texture display, Thunderbolt 5, or the extra GPU cores, holding an M3 Pro is still a strong position.

---All agents reported back! ├─ 🟠 Reddit: 3 threads │ 0 upvotes │ 0 comments ├─ 🔵 X: 12 posts │ 10 likes │ 1 repost ├─ 🔴 YouTube: 10 videos │ 7,866,702 views │ 3 with transcripts ├─ 🌐 Web: 0 pages (supplementary) └─ 🗣️ Top voices: @bhphoto (3 likes), @jameslmorton │ r/macbookpro

---I'm now an expert on M4 MacBook Pro reviews. Some things I can help with:

  • Compare the base M4 vs M4 Pro vs M4 Max for your specific workload (dev, creative, AI/ML)
  • Break down real-world battery life expectations by screen size and usage pattern
  • Advise whether upgrading from M1/M2/M3 is worth it based on long-term reviewer consensus

VERSION C

What I learned

Cross-referencing Reddit ownership reports, YouTube long-term reviews, and X commentary on the M4 MacBook Pro reveals a maturing consensus: this is excellent hardware with a surprisingly simple recommendation for most buyers. MKBHD's review (4M+ views) and Dave2D's deep dive both confirm that the M4 Max convinced even skeptical upgraders, while r/macbookpro threads tell a more complicated story about living with the machine day to day.

The Air-to-Pro switch is about more than the chip. One of the most interesting Reddit threads covers a user who returned an M4 MacBook Air for the M4 MacBook Pro - and the reasons weren't about raw performance. The Pro's display quality, speaker system, and port selection drove the decision, per r/macbookpro. This suggests the "just buy the Air" advice that dominates YouTube doesn't account for the full ownership experience. Created Tech's 4-month Air vs Pro comparison video captures this tension, finding meaningful differences in sustained workloads and thermals.

Battery expectations need recalibrating. Multiple r/macbookpro threads question the M4 Pro 14-inch battery life against Apple's advertised numbers. Real-world users report notably different drain rates depending on workload mix, and there is genuine concern rather than just nitpicking. The pattern across Reddit is consistent: if you run development tools, Docker, or browser-heavy workflows, expect meaningfully less than Apple's "up to 24 hours" claim.

Long-term reviewers are more candid. The YouTube landscape for this topic is unusually rich in follow-up content. Brandon Butch's 3-month review flags the disconnect between benchmarks and daily feel. Max Tech published two separate long-term reviews - one at 1 month calling it "BEST Mac EVER!?" and another at 6 months titled "Everyone was WRONG!" - showing how opinions evolve with extended use, per Max Tech on YouTube. Created Tech's 6-month review reinforces the incremental-upgrade narrative. Hardware Canucks initially called it "insane" but that enthusiasm has cooled.

Local AI is the sleeper use case. The X conversation is dominated by AI workload discussions. @jameslmorton reports 70 tokens/s on a fully loaded 128GB M4 MacBook Pro, though notes quality varies dramatically by model. @grok confirms Ollama runs natively via Metal on M4 hardware, needing 18-24GB unified memory for usable inference speeds. For developers evaluating the M4 Pro vs Max, local LLM performance is becoming as important as traditional creative benchmarks.

Configuration advice is converging. Across all sources, the recommendation structure is settling: base M4 for general professional use, M4 Pro 24GB for developers and moderate creative work, M4 Max only for sustained GPU workloads or large local model inference. MacRumors explicitly advises skipping Pro and Max for most buyers. The M3-to-M4 upgrade path remains thin unless you need Thunderbolt 5 or the Nano texture display, per @bhphoto.

---All agents reported back! ├─ 🟠 Reddit: 3 threads │ 0 upvotes │ 0 comments ├─ 🔵 X: 12 posts │ 10 likes │ 1 repost ├─ 🔴 YouTube: 10 videos │ 7,926,186 views │ 3 with transcripts ├─ 🟡 HN: 0 stories (no results this cycle) ├─ 🌐 Web: 0 pages (supplementary) └─ 🗣️ Top voices: @bhphoto (3 likes), @jameslmorton │ r/macbookpro

---I'm now an expert on M4 MacBook Pro reviews. Some things I can help with:

  • Compare the M4 MacBook Air vs Pro for your specific use case (the differences go beyond specs)
  • Recommend the right M4 configuration tier based on whether you run local AI models, creative apps, or dev tools
  • Give you realistic battery life expectations based on Reddit owner reports, not Apple's marketing numbers

SCORES

Version A

  • Groundedness: 4/5
  • Specificity: 4/5
  • Coverage: 4/5
  • Actionability: 4/5
  • Format: 3/5
  • Weighted Total: 3.85/5.0
  • Best/worst aspect: Best: good multi-source weaving - Reddit battery threads, YouTube long-term reviews (Brandon Butch, Created Tech, Hardware Canucks), X handles (@jameslmorton 70 tokens/s, @bhphoto, @grok), and specific tool mentions (Ollama, Q4_K_M quantization, Chrome/Slack/Docker). Worst: format has issues - the "---All agents" and "---I'm now" lines lack proper spacing, and HN shows 0 stories which is acknowledged but the narrative notes HN silence as a finding rather than just reporting the gap.

Version B

  • Groundedness: 4/5
  • Specificity: 4/5
  • Coverage: 3/5
  • Actionability: 4/5
  • Format: 3/5
  • Weighted Total: 3.70/5.0
  • Best/worst aspect: Best: solid specificity with named reviewers (Brandon Butch, Created Tech, Tech It Easy, Hardware Canucks, MacRumors, Just Josh), exact numbers (70 tokens/s, 128GB), and specific YouTube view data mentioned implicitly. Good narrative structure moving from configuration to battery to long-term to AI to upgrade advice. Worst: coverage is the weakest dimension - no HN data, and the synthesis doesn't weave X and YouTube together as tightly as it could. Stats block missing the HN line entirely.

Version C

  • Groundedness: 4/5
  • Specificity: 5/5
  • Coverage: 4/5
  • Actionability: 5/5
  • Format: 3/5
  • Weighted Total: 4.20/5.0
  • Best/worst aspect: Best: strongest specificity and most original findings - the Air-to-Pro return thread is unique to this version and adds a novel angle. Max Tech's dual review titles ("BEST Mac EVER!?" at 1 month vs "Everyone was WRONG!" at 6 months) perfectly illustrate opinion evolution. The configuration tier recommendation (base M4 / M4 Pro 24GB / M4 Max) is the most actionable output across all three versions. MKBHD (4M+ views) and Dave2D are explicitly named. Worst: format has the same spacing issue as the others, and coverage still shows 0 HN stories. YouTube view count (7.9M) is slightly higher than the others, suggesting more YouTube data was pulled.

VERDICT

Winner for M4 MacBook Pro review: Version C Why: Version C wins by surfacing unique angles the others miss (Air-to-Pro return story, Max Tech's evolving opinion across two reviews) and producing the most actionable configuration framework. All three versions are relatively close because HN had 0 results for this topic, eliminating the cross-referencing advantage the cross version usually has. Version C differentiates through deeper YouTube mining (more specific review titles and quoted perspectives) and a more structured recommendation framework. Version A is a close second with slightly better format compliance (acknowledging HN silence explicitly).

Reveal: {'A': 'hn', 'B': 'base', 'C': 'cross'}