Repository navigation
Tracking issue for RFC 2151, Raw Identifiers #48589
Description
Activity
- addedB-RFC-approvedBlocker: Approved by a merged RFC but not yet implemented.Blocker: Approved by a merged RFC but not yet implemented.T-langRelevant to the language teamRelevant to the language teamC-tracking-issueCategory: An issue tracking the progress of sth. like the implementation of an RFCCategory: An issue tracking the progress of sth. like the implementation of an RFC
on Feb 27, 2018 @Centril you rock! @petrochenkov, think you could supply some mentoring instructions here?
Reacted by Mazdak FarrokhzadI'd like to take a shot at this. It seems like it'd be a decent way to learn how rustc works.
Reacted by Esteban KuberThe relevant code is probably this, you'll want to make the
('r', Some('#'), _)case allow for the third character to be alphabetic or an underscore, and in that case skip the r and the#before runningident_continue.We could add an
is_rawboolean totoken::Identas well.You'll also need to feature gate this but we can do that later.
lmk if you have questions
Reacted by Rimi KanokawaWe could add an
is_rawboolean totoken::Identas well.This is possible, but would be unfortunate. Idents are used everywhere and supposed to be small.
Ideally we should limit the effect of
r#to lexer, for example by interningr#keywordidentifiers into separate slots (like gensyms) sor#keywordandkeywordhave differentNameSymbols.EDIT: The first paragraph is about
ast::Ident,token::Identmay actually be the appropriate place.Some clarifications are needed:
- How
r#affects context-dependent identifiers (aka weak keywords) likedefault.
Do they lose their special context-dependent properties and turn into "normal identifiers"?
// `union` is a normal ident, this is not an error union U { .... } // `union` is a raw ident, is this an error? r#union U { ... }
- How does
r#affect keywords that are "semantically special" and not "syntactically special"?
I'm talking about path segment keywords specifically.
For example,SelfinSelf::A::Bis already treated as normal identifier during parsing, it only gains special abilities during name resolution when we resolve identifiers namedSelf(orself/super/etc) in a special way.
#[derive(Default)] struct S; impl S { fn f() -> S { r#Self::default() // Is this an error? } }
- How
oh, I didn't realize we reuse Ident from the lexer.
I think
r#unionis an error (when used to create a union). We'll need the ident lexing step to return a bool on the lexed ident's raw-ness.I think it's ok for
r#Selfto work; but don't mind either wayAlso, lifetime identifiers weren't covered by the RFC -
r#'identor'r#ident.
(One more case of ident vs lifetime mismatch caused by lifetime token being a separate entity rather than a combination of'and identifier, cc https://internals.rust-lang.org/t/pre-rfc-splitting-lifetime-into-two-tokens/6716).I think it's fine if we don't have raw lifetime identifiers. Lifetimes are crate-local, their identifiers never need to be used by consumers of your crate, so lifetimes clashing with keywords can simply be fixed on epochs. Admittedly, writing a lint that makes that automatic may be tricky.
Raw identifiers are primarily necessary because people may need to call e.g. functions named
catch()in crates on an older epoch. This problem doesn't occur for lifetimes.Yeah, it's mostly a consistency question rather than a practical issue.
From what I've been seeing while looking around the codebase, I think the best way to implement this is to add a new parameter to
token::Ident, rather than messing with theSymbolitself?I think this would make implementing epoch-specific keywords easier, since there's no question of what
Symbolshould be used when, and in what epoch. (For example, you'd have to make sure theSymbolforcatchbeing used as an identifier in 2015 epoch code is the same as the Symbol forr#catchbeing used in a epoch where it's a full keyword.) This was already something I wasn't sure how to handle with contextual keywords.My main questions, right now, would be:
- Does this actually sound like the best approach?
- Reading the code, it looks like most feature gating is done on the AST after parsing, and not during parsing. Since nothing in the AST would reflect the raw identifiers being there at all with this approach, the feature check would have to be in
parser.rs. Would adding one there be an issue? How would I go about doing that, considering that module doesn't have other feature checks that I can use as a template? - A minor code style point: right now,
token::Identis declared asIdent(ast::Ident). To add anis_rawfield would mean having a mystery unnamedboolfield in a tuple struct, or making it use named fields, in which case, matching ontoken::Idents becomes nastier. One idea that did come to mind is adding anRawIdent(ast::Ident)variant, but then the compiler can't help me find places I might need to worry about raw identifiers. Any advice on this?
I'll implement lifetime parameters if it turns out to be easy to, I guess. As Manishearth said, it's not something you really need to escape ever.
Yes, we should not be affecting Symbol.
Regarding the feature gate, we can solve the problem later, but I was thinking of doing a delayed error or something since we don't know what feature gates are available whilst lexing
a mystery unnamed bool field in a tuple struct
I think that's fine. Folks usually do this as
Ident(ast::Ident, /* is_raw */ bool)I think it's best if we don't allow this to work for lifetime parameters, actually. We restrict the number of places where raw identifiers are allowed at a first pass, and if we need this for a later epoch, we add it then
60 remaining items
When it comes to new documentation, is there anything that should be updated besides the Reference? I'm not sure it belongs in the Book.
- added a commit that references this issue
on Aug 21, 2018 Merged! I think some boxes can be ticked off now, @Centril. :-)
As for the unresolved questions:
- Do macros need any special care with such identifier tokens?
No, I can't see how this would affect macros. Perhaps @petrochenkov could second this, however. - Should diagnostics use the r# syntax when printing identifiers that overlap keywords?
Yes, I think so, although possibly we should not do this if we're using an old edition and the keyword was only introduced in a later one. - Does rustdoc need to use the r# syntax? e.g. to document pub use old_epoch::*
I'm not sure of this. What do you mean bypub use old_epoch::*, @Centril?
- Do macros need any special care with such identifier tokens?
Does rustdoc need to use the r# syntax? e.g. to document pub use old_epoch::*
I'm not sure of this. What do you mean by pub use old_epoch::*For instance, if the
old_epochcrate had afn catch()which it could declare without bothering with raw identifiers, and anew_epochcrate used and re-exported it. The new crate would have had to use raw identifiers if it had declared that function itself. Should this affect the way rustdoc presents it?@cuviper Right, makes sense. I think the rules should be the same as what I proposed for diagnostics, in that case.
@Centril
I think it's still impossible for rust to reference external C++ symbols even with raw identifiers.
Because identifier names mangled by MS C++ ABI has characters such as '?', '@', which are not valid in rust identifiers.
For example, this is valid C++ code:#ifdef _MSC_VER // MS C++ ABI mangled identifier void* vtbl = __identifier("??_7BaseShader@pe3@@6B@"); #else // Itanium C++ ABI mangled name, all *nix toolchain use this. extern "C" void * _ZTVN3pe310BaseShaderE; void* vtbl = _ZTVN3pe310BaseShaderE ; #endifThe
__identifer("??_7BaseShader@pe3@@6B@")is a valid raw identifier in MS C++,
butr#??_7BaseShader@pe3@@6B@is not valid in rust.
Will raw identifier be relaxed to accept these characters?Inter-language-operability was not a goal of this RFC, it was purely limited to inter-edition-operability within Rust. AFAIK
#[link_name]should be enough for FFI purposes, there's no need to be able to have the identifiers in the Rust code identical to what is used to link to them.Reacted by scottmcmReacted by zephyr
This is a tracking issue for RFC 2151 (rust-lang/rfcs#2151).
Steps:
Unresolved questions:
Probably not.
r#syntax when printing identifiers that overlap keywords?Depends on the edition?
r#syntax? e.g. to documentpub use old_epoch::*