Newspaper3k
Unstable API
0.10.0
@project-lakechain/newspaper3k
The Newspaper3k middleware is based on the Newspaper3k library which provides an NLP model that is optimized for HTML article text extraction. It provides capability to analyze and extract the substance of HTML documents on the Web and use that text as an input to other middlewares in your pipeline.
📰 Extracting Text
Section titled “📰 Extracting Text”To use this middleware, you import it in your CDK stack and instantiate it as part of a pipeline.
import { Newspaper3kParser } from '@project-lakechain/newspaper3k';import { CacheStorage } from '@project-lakechain/core';
class Stack extends cdk.Stack { constructor(scope: cdk.Construct, id: string) { // The cache storage. const cache = new CacheStorage(this, 'Cache');
// The newspaper3k parser. const newspaper3k = new Newspaper3kParser.Builder() .withScope(this) .withIdentifier('Newspaper3k') .withCacheStorage(cache) .withSource(source) // 👈 Specify a data source .build(); }}🏗️ Architecture
Section titled “🏗️ Architecture”This middleware is based on a Lambda compute running the Newspaper3k library packaged as a Docker container.

🏷️ Properties
Section titled “🏷️ Properties”Supported Inputs
Section titled “Supported Inputs”| Mime Type | Description |
|---|---|
text/html |
HTML documents. |
Supported Outputs
Section titled “Supported Outputs”| Mime Type | Description |
|---|---|
text/plain |
Plain text documents. |
Supported Compute Types
Section titled “Supported Compute Types”| Type | Description |
|---|---|
CPU |
This middleware only supports CPU compute. |
📖 Examples
Section titled “📖 Examples”- Article Curation Pipeline - Builds a pipeline converting HTML articles into plain text.