Recently, Facebook parent company Meta released a version that, to be honest, is indeed a bit too similar to Google’s version in terms of various functions – both can convert PDF files into a two-person dialogue transcript similar to a Podcast, and then use a text-to-speech model to generate Podcast-like audio files for subsequent use. Continue reading: Meta is also launching NotebookLlama, which can generate “Podcast-like” content; how practical is it? See the article for details.

▲Image source: Meta
Meta also launches NotebookLlama, which can generate “podcast-like” content—check this article to see how practical it really is.
Although both the name and the description on Github make it clear that Meta’s NotebookLlama is paying homage to Google’s NotebookLM — especially since the description directly calls it “NotebookLlama: An Open Source version of NotebookLM.”

Speaking of it, this version recently released by Meta, Facebook’s parent company, is frankly a bit too similar to Google’s version in terms of various features – both can convert PDF files into a two-person dialogue script similar to a podcast, and then use a speech-to-text model to generate a podcast-like audio file for subsequent applications.
Because of this, the point people are likely more concerned about is, beyond the differences in open source, how Meta NotebookLlama actually performs. Below is its initial actual dialogue and text-writing performance as posted online:
Wow! Meta dropped an open NotebookLM recipe: NotebookLlama 🔥
It uses L3.2 1B/ 3B for pre-processing the PDF, L3.1 70B for Transcript creation, L3.1 8B for re-writes and Parler TTS for Text to Speech ⚡
Step 1: Pre-process PDF: Use Llama-3.2-1B-Instruct to pre-process the PDF… pic.twitter.com/L7hb5GsMtl
— Vaibhav (VB) Srivastav (@reach_vb) October 27, 2024
Okay, after hearing its performance, most people would probably think that compared to Google NotebookLM, Meta’s version still needs work – well, it still sounds pretty robotic, and the conversational break points are a bit off.
In fact, the officials have also stated that the current version is mainly limited by the model’s performance, which is why it doesn’t feel natural enough. However, they also put forward a similar argument to that seen in the early development of many models — or, in some cases, throughout their entire development — namely, that it has the potential to achieve better progress through subsequent model changes or improvements.

At this stage, NotebookLlama’s workflow involves converting files such as PDFs into TXT files. Next, after transforming into a “podcast-like” version, the content is re-polished through a model to give it more dramatic tension—or rather, a performative style. Finally, the text-to-speech part, which is blamed for the current voice still not sounding natural enough, will focus on enhancing the conversational feel.
Of course, this series of models basically all involve Llama’s large language models (with the speech being parler-tts/parler-tts-mini-v1 and bark/suno). I’m also personally quite curious whether, with more validation through real-world usage in the future, it can demonstrate distinct advantages compared to NotebookLM.
Citation source:GithubVia:TechCrunch|
Further Reading:
Regarding the platform’s video-quality reduction policy that 影視颶風 called out, Instagram has its own take: it instead reduces quality for low-view-count content.
Source: KOCPC Chinese