The processing pipeline
When you upload a document to Upstream, it goes through a multi-step pipeline that turns raw text into structured, searchable knowledge. Here's what happens at each stage.
1. Upload
Your document is received and added to the processing queue. You'll see it appear on the Upload page with a status indicator.
2. Parse
Upstream extracts the text content from your document. Supported file types include:
- DOCX (Microsoft Word)
- TXT (plain text)
- And other common document formats
The parser handles formatting, tables, and multi-page documents automatically.
3. Speaker identification
For transcripts (such as sales calls, meetings, or interviews), Upstream identifies individual speakers so it can attribute quotes and context to the right person.
4. Entity extraction
Upstream analyzes your text and identifies key entities — Pain Points, Objections, Feature Requests, and 7 other types. Each entity receives a confidence score from 0 to 100 based on how clearly it appears in your document.
5. Relationship mapping
Connections between entities are discovered automatically. For example, Upstream might identify that a Pain Point leads_to a Feature Request, or that Customer Feedback supports a Market Insight.
6. Embedding generation
Your content is indexed for semantic search. This means when you ask questions in chat, Upstream can find relevant information even if you use different words than the original document.
7. Review
Entities with high confidence scores (80 and above) are auto-approved and immediately enter your knowledge graph. Entities with lower confidence scores are sent to your Review Queue, where you can approve, edit, or reject them.
Check processing status
Visit the Upload page to see the status of all your documents. Each document shows its current processing stage so you always know where things stand.
What to do if processing seems stuck
Processing usually completes within a few minutes. If a document appears stuck:
- Wait a few more minutes. Larger documents and complex PDFs take longer to process.
- Refresh the page. The status may have updated.
- Re-upload the document. If the status hasn't changed after 15 minutes, try uploading the document again.
- Check the file. Make sure the document isn't corrupted, password-protected, or an unsupported format.
Tips for documents that extract well
Not all documents are created equal. You'll get better results with:
- Clear structure. Documents with headings, bullet points, and short paragraphs produce more accurate extractions.
- Specific language. Concrete statements ("Customers say pricing is confusing") extract better than vague ones ("There are some concerns").
- Transcripts with speaker labels. If you're uploading a transcript, include speaker names for better attribution.
- Focused content. A single-topic document often extracts better than a long, multi-topic one. Consider splitting large documents into smaller, focused files.
Related articles
Still need help?
Email us at hello@getupstream.ai — we read every message.