Howell: analysis of blog comments
Table of Contents
Introduction
Jacobovici's "Exodus proofs" video provides the only detailed blog analysis of scale that I have done (as of 18Oct2024). While I found the video to be [excellent, stunning, rich with details new to me], in the end it has been the blog analysis of 22,000+ comments by youTube viewers that provides a [deep, rich, fascinating] opportunity to look into some aspects of [what, how, why] people [think, believe]. That in turn affords me a strong basis for questioning my own [think, believe]s. Not apparent from this analysis are the even deeper subjects listed in the section ?? below:
Transcription notes :
+-----+
-->
- The video text is SUPPOSED to be a literal transcription of the video commentary. The vast bulk of the transcription is done by Google's "AI machines", as is sometimes the case for some youTube videos. However, the machine output has some [error, omission]s, and it lacks even minimal formatting.
- Perhaps more important that machine mis-transcriptions, is the failed clarity of human [pro, e]nunciation of [word, phrase, sentence]s, including [stutter, wrong word, backtrack, mis-construct]s. I have found that in Google transcriptions of [my own, friends'] voice recordings as well. So let's NOT blame everything on the machines, just to protect our own pathetic butts!
- I have tried to guess at wording where the machine transcription from audio wasn't clear to me. So far, I haven't even tried to listen through the video to make corrections, as it's easiest to clean up the text formatting first to have something to work with.
- Partially restructured for easier reading, timestamps are a bit off (easier for me).
- I haven't corrected [name, place, spelling, etc]s, many names must be tryped in from captions (later?...)
- <machine transcription text = ?Howell's guess?> denote possible mis-wordings of the transcription. These are retained, as I typically do some tracking of Google translate performance to help predict when it is wothwhile, and when simply hand typing voice audio may be preferred.
- The machine transcription (Large Language Model?)is better than typing everything by hand, but it is still Kludgy.
Tools to [download, analyse, etc] * [webPages, blogs, speech]
The issue of screening through oceans of [emotion, troll, lazy, delinquent, hypocritical, dishonest, manipulative] comments is a fun idea.
For me, the [download, AI analysis] is HUGE, as there are typically [rare, brilliant] comments in blogs, but possibly less than one-in-1,000. Anything that helps find the gold is almost worth gold. This can turn blogs from cesspools into gold mines. Not to mention that my writing of scripts gets to be a huge time consumer, and is much less effective.
As a "workable" text, powerful Unix tools can be used to [search, analyze, change, extract] information on this webPage. The same tools can be used for analysis across many [txt, html, etc] files, including this webPage. But it's not as fun as playing with the new toys above. Until they don't do what I need, then I'll build something myself (cheap and simple, but very powerful).
-
Google transcript text-to-speech - youTube video's, such as Jacovici's Exodus proofs" often provide the transcription. I subscribe to Google's transcription service for my work, and I want to use it far more often when I clear out a backlog of other. Great for taking voice notes while traveling etc, capturing reminiscences of older [relative, friend]s.
-
Bartek Igielski's youTube-comments-downloader - Scraping blogs to grab comments is a pain for me, and this tool was [fast, great]. I purchased this (129$US). The html download of the Jacobovici video comments was 58.7 Mb (large), and the txt file was 6.2 Mb. Unfortunately, I didn't see any mention beyond sentiment analysis etc, and I won't pay for a chatGPT upgrade, for now.
-
Bartek Igielski's comments analyzer built on openAI's chatGPT. Actually, I didn't upgrade to chatGPT+ ($), because it seemed that the announced capabilties weren't what I was looking for.
-
Howell's example bash scripts
Howell's blog analysis problems:
- Basic problem is that 22,0000+ comments are too many to read.
I needed to get this down to ~100-200.
- If I annotated [Inn, Out] again, I would get very different results. Arbitrary.
- Many comments on the same theme, but I tried to keep only a few.
- I lost many classification by improper file saving. Much work lost, may have lost concepts
when re-doing, as I remember similar comments, but were they saved?
- [Inn, Out] classification required reading ALL pCmtScr. This was very tiring, and the
quality of my work declined. I tended to exclude more of the comments.
- some criteria for comment EXclusion :
- same [detail, concept]s as provided in other comments. Which of several
similar comments is retained was an arbitrary decision (not fair to excluded bloggers).
- repeated comments (sometimes repeated 4 or 5 times!), only 1st retained, if pertinent
- [derogatory, ad hominen] * [phrase, cynicism]s
- mostly [affirmation, rejection] of [faith, religion]
- weird [purity, restrictiveness] of what [science, religion] does
- some criteria for comment INclusion :
- [keyWord, concept]s directly related to details in the video, or alternate to it
- [relevant, specific] * [author, reference]s are mentioned
- [novel, unique]*[concept, analysis, detail] with repect to all other comments
(this is my guess, as I certainly haven't double-checked against 22,000 comments.)
- [interesting, broad] descriptions of [history, science, religion] context
- focus on motives beyond the film coverage [greed, gold, evil, etc].
Maybe I'm just tired of cospiracy theor[y, ists]... and prefer to [hear, read] people who have more of a basis.
-
-
-
-
-
-
-
-
-
- the above criteria were [loose, inconsistent]ly applied
-
-
-
-
- The above [IN, EX]clusion criteria were [loose, inconsistent]ly applied, and depended on remembering what similar detail I had already seen. (My memory isn't so good, especially with age.)
Blogs: [deep, broad, high]er themes
[rapid, complete] flips of people's thinking (eg in science)
People's long-standing beliefs can change completely even 180 degrees, or to completely different [basis, dimensionality] without them being aware of it. If they are aware of it, the usual explanation is that breakthrough data was the cause, even though that may have been around for [year, decade, century, millenia]s.
Blog explosion of [diversity, dimensionality]
Blogs can greatly increase the [richness, context] of a [video, webPage, book], but are likely under-utilised for that purpose. A primary challenge is to find the specs of gold (according to your interests) in a hill of [dirt, gravel]. My amateur guess is that :
- most bloggers just need to "say something that hey are thinking about", just as most normal conversations seem to be driven by an urg to talk (blabber-driven). Not surprisingly, blog comments are often un-related to the [video, webPage, book], and many comments are pure emotion-speaking (eg "it's a nice day, 22 Celsius", "that politician is a worm", etc).
-
-
-
-
-
-
-
The [blessing, curse] of dimensionality
the [blessing, curse] of high dimensionality. Alexandre Gorban's video is a great explanation of long-standing awareness that the [statistic, causitive] nature of very high dimensional systems is in stark contract, even opposite, to conventional systems. Few people formally address that, although we probably all handle that "naturally" to some extent?
Consensus is poison. Politics needs this poison.
"Consensus can be poison" to thinking, even though it may be essential for politics? I made mistakes in my blog analysis by excluding whole areas of blog comments that weren't directly related to the [data, analysis] of the video (eg [Bible, archaeology, geology, astronomy, etc]), but it would be an [open, huge, never-ending] quest in practice. Blog comments only cover a tiny fraction of relevant thematic dimensions.
The [hypothesis, theory] should stand on its own in the face of [independent, 3rd party]
[motherhood, day car, kindergarten] constraints of modern intellectualism
Programming the programmers
Conscious mind, resonant brain
Stephen Grossberg's book "Conscious mind, resonant brain" is the best reference that I know of that really brings out a solid basis for [cooperative-competetive, top down - bottom up, resonant resolving of a likely reality from very messy basic perceptions, Adaptive Resonance Theory (ART), laminar computing, consciousness, etc]. This is a whole new world for me.
Cooperative-competitive, top down - bottom-up
Progress stops because of the most [decorated, successful] scientific [theory, expert]s
Amateurs versus experts
Amateurs have always been inventors of new [concept, theory]s, but one rarely sees them being credit. Here I include professional scientists that are not experts in a scientific domain, or often-enough who are employees in [small, non-Western] univrsities. After all, they are openly disregard by experts. Scientists often steal these ideas, stab the amateurs in the back, and manipulate credit to themselves.
The internet, and open publishing, has greatly enhanced the accessibility of amateurs to journal articles, and greatly increased their impact, albeit rarely recognized by "Official science" (i.e. [government, academic] scientists).
Before double-dipping experts benefitted from taxpayer investments.
(not hiding behind a paywall so that only go to
Amateurs permit one to most easily avoid the turbic thinking of essentially all scientists, and especially the most [successful, awarded, famous] scientists.
[nonsense, science, nonscience]
Sites that profile scientific fraud: new one that had Ben Davidon book, Journal of irreproducible results, etc
Experts often impede progress: only able to focus on one "truth" at a time, fail to understand limitations of data, can't change thinking fast enough with new [data, concepts]; must not risk membership in "their club of like-minded"; not strong enough to stand on their own: must follow the leader; and the leader must have ready-made followers
???
???