The Brain Language Metrics on Company Filings (BLMCF) dataset has the objective of monitoring several language metrics on 10-Ks and 10-Qs company reports for approximately 6000+ US stocks. Example of metrics are financial sentiment, percentage of specific language type in the document (e.g. litigious language) and similarity among documents. This extended version provides additional language metrics and an analysis of the whole report together with specific report sections (e.g. Risk Factors).
The Brain Language Metrics on Company Filings (BLMCF) dataset has the objective of monitoring several language metrics on 10-Ks and 10-Qs company reports for approximately 6000+ US stocks.
Some literature works claim inefficiencies in the market response to company filings information due to the increased complexity and length of such reports; over the last 20 years, the length of the average 10-K has in fact increased dramatically.
Our dataset is made of two parts; the first one includes the language metrics of the most recent 10-K or 10-Q report for each firm, namely:
Financial sentiment
Percentage of words belonging to financial domain classified by language types: constraining, litigious, uncertainty and interesting language.
Readability score
Lexical metrics such as lexical density and richness
Text statistics such as the report length and the average sentence length
The second part includes the differences between the two most recent 10-Ks or 10-Qs reports of the same period for each company, namely:
Difference of the various language metrics (e.g. delta sentiment, delta readability score delta, delta percentage of a specific language type etc.)
Similarity metrics between documents, also with respect to a specific language type (for example similarity with respect to “litigious” language or “uncertainty” language)
Our dataset includes the metrics and related differences both for the whole report and for specific sections (Risk Factors and Management Discussion and Analysis)
Feed Details
The dataset is updated with a daily frequency since new 10-Ks and 10-Qs reports are released every day for some of the universe companies. Clearly the largest update will be around February, April, August and November when the largest number of reports is released. The historical dataset is available from year 2010.
The dataset contains historical data from January 2010 that can be freely accessed for 2 months. For a live feed please contact us at support@braincompany.co and we will make accessible a customized version of the product on AWS Data Exchange according to Client requirements.
Disclaimer
The content of this dataset is not to be intended as investment advice. The material is provided for informational purposes only and does not constitute an offer to sell, a solicitation to buy, or a recommendation or endorsement for any security or strategy, nor does it constitute an offer to provvaluee investment advisory or other services by Brain. Brain makes no guarantees regarding the accuracy and completeness of the information expressed in the dataset.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
This is a free trial listing with a single pricing dimension. Product Access (Units) grants you access to the dataset at no charge. There are no tiers, usage add-ons, or scaling options to compare. You get one flat access dimension, which lets you explore the historical language-metrics dataset covering 10-K and 10-Q filings for US stocks. Because it is a trial with one dimension, pricing does not change based on usage volume or quantity.
Top-of-mind questions for buyers
What does the Product Access unit give me access to in this trial?
It grants access to the historical language-metrics dataset. The data covers 10-K and 10-Q filings for 6000+ US stocks. Metrics include financial sentiment, readability scores, lexical density and richness, similarity measures, and differences between documents. The historical dataset reaches back to 2010.
How often is the dataset updated during the trial?
The dataset updates daily, since new filings are released regularly for companies in the universe. The largest updates fall around February, April, August, and November, when the most reports are filed. Your single access dimension covers this ongoing update flow at no charge.
braincompany.co
Helpful?
Vendor refund policy
No refunds are offered for this product, for more information please contact support@braincompany.co
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
The Brain Language Metrics on Earnings Calls Transcripts (BLMECT) dataset has the objective of monitoring several language metrics for the quarterly earnings call transcripts of 4500+ US stocks.
With this dataset we aim at providing additional building blocks to asset managers to build investment strategies based on alternative data.
The Brain Sentiment Indicator monitors public financial news for 6000+ stocks from about 2000 financial media sources in 33 languages by measuring financial sentiment, number of stories published and buzz.
The sentiment scoring technology is based on a combination of various natural language processing techniques.
For each stock the sentiment score corresponds to the average of sentiment for each piece of news and it is available on two time scales; 7 days and 30 days.
The Brain Sentiment Indicator (Crypto version) monitors public financial news for 60+ of the major cryptocurrencies from thousands of financial media sources in 33 languages. The sentiment scoring technology is based on a combination of various Natural Language Processing techniques.
The sentiment score assigned to each crytpo is a value ranging from -1 (most negative) to +1 (most positive) that is updated with a daily frequency.
Brain Machine Learning proprietary platform is exploited to generate a daily stock ranking based on the predicted future returns of a universe of 1000 stocks on five time horizons: 2,3, 5, 10 and 21 days (other time horizons could be developed and tested upon request). The model implements specific machine learning techniques to combine a variety of features with a series of techniques aimed at mitigating the well-known overfitting problem for financial data with a low signal to noise ratio.